voice activity detector Latest Research Papers

Voice recognition plays a key role in spoken communication that helps to identify the emotions of a person that reflects in the voice. Gender classification through speech is a widely used Human Computer Interaction (HCI) as it is not easy to identify gender by computer. This led to the development of a model for “Voice feature extraction for Emotion and Gender Recognition”. The speech signal consists of semantic information, speaker information (gender, age, emotional state), accompanied by noise. Females and males have different voice characteristics due to their acoustical and perceptual differences along with a variety of emotions which convey their own unique perceptions. In order to explore this area, feature extraction requires pre- processing of data, which is necessary for increasing the accuracy. The proposed model follows steps such as data extraction, pre- processing using Voice Activity Detector (VAD), feature extraction using Mel-Frequency Cepstral Coefficient (MFCC), feature reduction by Principal Component Analysis (PCA) and Support Vector Machine (SVM) classifier. The proposed combination of techniques produced better results which can be useful in the healthcare sector, virtual assistants, security purposes and other fields related to the Human Machine Interaction domain.

Download Full-text

A 28.5µW All-Analog Voice-Activity Detector

2021 IEEE International Symposium on Circuits and Systems (ISCAS) ◽

10.1109/iscas51556.2021.9401504 ◽

2021 ◽

Author(s):

Udita Mukherjee ◽

Tanmay Halder ◽

Anand Kannan ◽

Sovan Ghosh ◽

Shanthi Pavan

Keyword(s):

Voice Activity Detector ◽

Voice Activity

Download Full-text

Voice Feature Extraction for Gender and Emotion Recognition

ITM Web of Conferences ◽

10.1051/itmconf/20214003008 ◽

2021 ◽

Vol 40 ◽

pp. 03008

Author(s):

Madhu M. Nashipudimath ◽

Pooja Pillai ◽

Anupama Subramanian ◽

Vani Nair ◽

Sarah Khalife

Keyword(s):

Feature Extraction ◽

Data Extraction ◽

Principal Component ◽

Feature Reduction ◽

Healthcare Sector ◽

Support Vector ◽

Svm Classifier ◽

Human Machine Interaction ◽

Interaction Domain ◽

Voice Activity Detector

Voice recognition plays a key function in spoken communication that facilitates identifying the emotions of a person that reflects within the voice. Gender classification through speech is a popular Human Computer Interaction (HCI) method on account that determining gender through computer is hard. This led to the development of a model for "Voice feature extraction for Emotion and Gender Recognition". The speech signal consists of semantic information, speaker information (gender, age, emotional state), accompanied by noise. Females and males have specific vocal traits because of their acoustical and perceptual variations along with a variety of emotions which bring their own specific perceptions. In order to explore this area, feature extraction requires pre-processing of data, which is necessary for increasing the accuracy. The proposed model follows steps such as data extraction, pre-processing using Voice Activity Detector(VAD), feature extraction using Mel-Frequency Cepstral Coefficient(MFCC), feature reduction by Principal Component Analysis(PCA) and Support Vector Machine (SVM) classifier. The proposed combination of techniques produced better results which can be useful in healthcare sector, virtual assistants, security purposes and other fields related to Human Machine Interaction domain.

Download Full-text

A new joint noise reduction and echo suppression system based on FBSS and automatic voice activity detector

Applied Acoustics ◽

10.1016/j.apacoust.2020.107444 ◽

2020 ◽

Vol 168 ◽

pp. 107444

Author(s):

Rahima Henni ◽

Mustapha Djebari ◽

Mohamed Djendi

Keyword(s):

Noise Reduction ◽

Voice Activity Detector ◽

Echo Suppression ◽

Voice Activity ◽

Suppression System

Download Full-text

Online Speech Recognition Using Multichannel Parallel Acoustic Score Computation and Deep Neural Network (DNN)- Based Voice-Activity Detector

Applied Sciences ◽

10.3390/app10124091 ◽

2020 ◽

Vol 10 (12) ◽

pp. 4091 ◽

Cited By ~ 1

Author(s):

Yoo Rhee Oh ◽

Kiyoung Park ◽

Jeon Gyu Park

Keyword(s):

Neural Network ◽

Speech Recognition ◽

Deep Neural Network ◽

Computation Method ◽

Acoustic Model ◽

Voice Activity Detector ◽

Context Sensitive ◽

Speech Features ◽

Voice Activity ◽

Main Thread

This paper aims to design an online, low-latency, and high-performance speech recognition system using a bidirectional long short-term memory (BLSTM) acoustic model. To achieve this, we adopt a server-client model and a context-sensitive-chunk-based approach. The speech recognition server manages a main thread and a decoder thread for each client and one worker thread. The main thread communicates with the connected client, extracts speech features, and buffers the features. The decoder thread performs speech recognition, including the proposed multichannel parallel acoustic score computation of a BLSTM acoustic model, the proposed deep neural network-based voice activity detector, and Viterbi decoding. The proposed acoustic score computation method estimates the acoustic scores of a context-sensitive-chunk BLSTM acoustic model for the batched speech features from concurrent clients, using the worker thread. The proposed deep neural network-based voice activity detector detects short pauses in the utterance to reduce response latency, while the user utters long sentences. From the experiments of Korean speech recognition, the number of concurrent clients is increased from 22 to 44 using the proposed acoustic score computation. When combined with the frame skipping method, the number is further increased up to 59 clients with a small accuracy degradation. Moreover, the average user-perceived latency is reduced from 11.71 s to 3.09–5.41 s by using the proposed deep neural network-based voice activity detector.

Download Full-text