Automatic Speech Recognition of Continuous Speech Signal of Gujarati Language Using Machine Learning

Stuttering or Stammering is a speech defect within which sounds, syllables, or words are rehashed or delayed, disrupting the traditional flow of speech. Stuttering can make it hard to speak with other individuals, which regularly have an effect on an individual's quality of life. Automatic Speech Recognition (ASR) system is a technology that converts audio speech signal into corresponding text. Presently ASR systems play a major role in controlling or providing inputs to the various applications. Such an ASR system and Machine Translation Application suffers a lot due to stuttering (speech dysfluency). Dysfluencies will affect the phrase consciousness accuracy of an ASR, with the aid of increasing word addition, substitution and dismissal rates. In this work we focused on detecting and removing the prolongation, silent pauses and repetition to generate proper text sequence for the given stuttered speech signal. The stuttered speech recognition consists of two stages namely classification using LSTM and testing in ASR. The major phases of classification system are Re-sampling, Segmentation, Pre-Emphasis, Epoch Extraction and Classification. The current work is carried out in UCLASS Stuttering dataset using MATLAB with 4% to 6% increase in accuracy when compare with ANN and SVM.

Download Full-text

Introduction to Speech Recognition

Intelligent Information Technologies ◽

10.4018/978-1-59904-941-0.ch007 ◽

2011 ◽

pp. 141-161

Author(s):

Sergio Suárez-Guerra ◽

Jose Luis Oropeza-Rodriguez

Keyword(s):

Artificial Intelligence ◽

Signal Processing ◽

Speech Recognition ◽

Automatic Speech Recognition ◽

Speech Signal ◽

Essential Information ◽

Science Field ◽

Applied Artificial Intelligence ◽

And Linguistics ◽

Successful Technology

This chapter presents the state-of-the-art automatic speech recognition (ASR) technology, which is a very successful technology in the computer science field, related to multiple disciplines such as the signal processing and analysis, mathematical statistics, applied artificial intelligence and linguistics, and so forth. The unit of essential information used to characterize the speech signal in the most widely used ASR systems is the phoneme. However, recently several researchers have questioned this representation and demonstrated the limitations of the phonemes, suggesting that ASR with better performance can be developed replacing the phoneme by triphones and syllables as the unit of essential information used to characterize the speech signal. This chapter presents an overview of the most successful techniques used in ASR systems together with some recently proposed ASR systems that intend to improve the characteristics of conventional ASR systems.

Download Full-text

Introduction to Speech Recognition

Advances in Audio and Speech Signal Processing ◽

10.4018/978-1-59904-132-2.ch011 ◽

2011 ◽

pp. 325-348 ◽

Cited By ~ 2

Author(s):

Sergio Suárez-Guerra ◽

Jose Luis Oropeza-Rodriguez

Keyword(s):

Artificial Intelligence ◽

Signal Processing ◽

Speech Recognition ◽

Automatic Speech Recognition ◽

Speech Signal ◽

Essential Information ◽

Science Field ◽

Applied Artificial Intelligence ◽

And Linguistics ◽

Successful Technology

This chapter presents the state-of-the-art automatic speech recognition (ASR) technology, which is a very successful technology in the computer science field, related to multiple disciplines such as the signal processing and analysis, mathematical statistics, applied artificial intelligence and linguistics, and so forth. The unit of essential information used to characterize the speech signal in the most widely used ASR systems is the phoneme. However, recently several researchers have questioned this representation and demonstrated the limitations of the phonemes, suggesting that ASR with better performance can be developed replacing the phoneme by triphones and syllables as the unit of essential information used to characterize the speech signal. This chapter presents an overview of the most successful techniques used in ASR systems together with some recently proposed ASR systems that intend to improve the characteristics of conventional ASR systems.

Download Full-text

Machine Learning in Automatic Speech Recognition: A Survey

IETE Technical Review ◽

10.1080/02564602.2015.1010611 ◽

2015 ◽

Vol 32 (4) ◽

pp. 240-251 ◽

Cited By ~ 28

Author(s):

Jayashree Padmanabhan ◽

Melvin Jose Johnson Premkumar

Keyword(s):

Machine Learning ◽

Speech Recognition ◽

Automatic Speech Recognition

Download Full-text

Towards a continuous speech corpus for banking domain automatic speech recognition

2017 International Conference on Speech Technology and Human-Computer Dialogue (SpeD) ◽

10.1109/sped.2017.7990436 ◽

2017 ◽

Cited By ~ 1

Author(s):

George Suciu ◽

Stefan-Adrian Toma ◽

Romulus Cheveresan

Keyword(s):

Speech Recognition ◽

Automatic Speech Recognition ◽

Continuous Speech ◽

Speech Corpus

Download Full-text

Automatic speech signal segmentation based on the innovation adaptive filter

International Journal of Applied Mathematics and Computer Science ◽

10.2478/amcs-2014-0019 ◽

2014 ◽

Vol 24 (2) ◽

pp. 259-270 ◽

Cited By ~ 9

Author(s):

Ryszard Makowski ◽

Robert Hossa

Keyword(s):

Speech Recognition ◽

Automatic Speech Recognition ◽

Adaptive Filter ◽

Speech Signal ◽

Detection Efficiency ◽

Difficult Problem ◽

Second Order ◽

Signal Segmentation ◽

Schur Algorithm ◽

Starting Point

Abstract Speech segmentation is an essential stage in designing automatic speech recognition systems and one can ﬁnd several algorithms proposed in the literature. It is a difﬁcult problem, as speech is immensely variable. The aim of the authors’ studies was to design an algorithm that could be employed at the stage of automatic speech recognition. This would make it possible to avoid some problems related to speech signal parametrization. Posing the problem in such a way requires the algorithm to be capable of working in real time. The only such algorithm was proposed by Tyagi et al., (2006), and it is a modiﬁed version of Brandt’s algorithm. The article presents a new algorithm for unsupervised automatic speech signal segmentation. It performs segmentation without access to information about the phonetic content of the utterances, relying exclusively on second-order statistics of a speech signal. The starting point for the proposed method is time-varying Schur coefﬁcients of an innovation adaptive ﬁlter. The Schur algorithm is known to be fast, precise, stable and capable of rapidly tracking changes in second order signal statistics. A transfer from one phoneme to another in the speech signal always indicates a change in signal statistics caused by vocal track changes. In order to allow for the properties of human hearing, detection of inter-phoneme boundaries is performed based on statistics deﬁned on the mel spectrum determined from the reﬂection coefﬁcients. The paper presents the structure of the algorithm, deﬁnes its properties, lists parameter values, describes detection efﬁciency results, and compares them with those for another algorithm. The obtained segmentation results, are satisfactory.

Download Full-text

Voice-Based Speaker Identification and Verification

Advances in Library and Information Science - Handbook of Research on Knowledge and Organization Systems in Library and Information Science ◽

10.4018/978-1-7998-7258-0.ch016 ◽

2021 ◽

pp. 288-316

Author(s):

Keshav Sinha ◽

Rasha Subhi Hameed ◽

Partha Paul ◽

Karan Pratap Singh

Keyword(s):

Speech Recognition ◽

Automatic Speech Recognition ◽

Speech Signal ◽

Reference Model ◽

Speaker Identification ◽

Recognition System ◽

Speech Recognition System ◽

Primary Focus ◽

Dynamic Time ◽

Dynamic Time Wrapping

In recent years, the advancement in voice-based authentication leads in the field of numerous forensic voice authentication technology. For verification, the speech reference model is collected from various open-source clusters. In this chapter, the primary focus is on automatic speech recognition (ASR) technique which stores and retrieves the data and processes them in a scalable manner. There are the various conventional techniques for speech recognition such as BWT, SVD, and MFCC, but for automatic speech recognition, the efficiency of these conventional recognition techniques degrade. So, to overcome this problem, the authors propose a speech recognition system using E-SVD, D3-MFCC, and dynamic time wrapping (DTW). The speech signal captures its important qualities while discarding the unimportant and distracting features using D3-MFCC.

Download Full-text

Continuous Speech Recognition of Kazakh Language

ITM Web of Conferences ◽

10.1051/itmconf/20192401012 ◽

2019 ◽

Vol 24 ◽

pp. 01012 ◽

Cited By ~ 2

Author(s):

Оrken Mamyrbayev ◽

Mussa Turdalyuly ◽

Nurbapa Mekebayev ◽

Kuralay Mukhsina ◽

Alimukhan Keylan ◽

...

Keyword(s):

Neural Networks ◽

Speech Recognition ◽

Error Rate ◽

Speech Signal ◽

Deep Neural Networks ◽

Continuous Speech ◽

Continuous Speech Recognition ◽

Reliable System ◽

Word Error Rate ◽

Language Studies

This article describes the methods of creating a system of recognizing the continuous speech of Kazakh language. Studies on recognition of Kazakh speech in comparison with other languages began relatively recently, that is after obtaining independence of the country, and belongs to low resource languages. A large amount of data is required to create a reliable system and evaluate it accurately. A database has been created for the Kazakh language, consisting of a speech signal and corresponding transcriptions. The continuous speech has been composed of 200 speakers of different genders and ages, and the pronunciation vocabulary of the selected language. Traditional models and deep neural networks have been used to train the system. As a result, a word error rate (WER) of 30.01% has been obtained.

Download Full-text