Deep Learning Approaches for Speech Emotion Recognition

Author(s):  
Anjali Bhavan ◽  
Mohit Sharma ◽  
Mehak Piplani ◽  
Pankaj Chauhan ◽  
Hitkul ◽  
...  
2021 ◽  
Vol 4 (4) ◽  
pp. 273
Author(s):  
Nhat Truong Pham ◽  
Ngoc Minh Duc Dang ◽  
Sy Dung Nguyen

Feature extraction and emotional classification are significant roles in speech emotion recognition. It is hard to extract and select the optimal features, researchers can not be sure what the features should be. With deep learning approaches, features could be extracted by using hierarchical abstraction layers, but it requires high computational resources and a large number of data. In this article, we choose static, differential, and acceleration coefficients of log Mel-spectrogram as inputs for the deep learning model. To avoid performance degradation, we also add a skip connection with dilated convolution network integration. All representatives are fed into a self-attention mechanism with bidirectional recurrent neural networks to learn long term global features and exploit context for each time step. Finally, we investigate contrastive center loss with softmax loss as loss function to improve the accuracy of emotion recognition. For validating robustness and effectiveness, we tested the proposed method on the Emo-DB and ERC2019 datasets. Experimental results show that the performance of the proposed method is strongly comparable with the existing state-of-the-art methods on the Emo-DB and ERC2019 with 88% and 67%, respectively. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium provided the original work is properly cited.


Sensors ◽  
2021 ◽  
Vol 21 (4) ◽  
pp. 1249
Author(s):  
Babak Joze Abbaschian ◽  
Daniel Sierra-Sosa ◽  
Adel Elmaghraby

The advancements in neural networks and the on-demand need for accurate and near real-time Speech Emotion Recognition (SER) in human–computer interactions make it mandatory to compare available methods and databases in SER to achieve feasible solutions and a firmer understanding of this open-ended problem. The current study reviews deep learning approaches for SER with available datasets, followed by conventional machine learning techniques for speech emotion recognition. Ultimately, we present a multi-aspect comparison between practical neural network approaches in speech emotion recognition. The goal of this study is to provide a survey of the field of discrete speech emotion recognition.


Author(s):  
Hanina Nuralifa Zahra ◽  
Muhammad Okky Ibrohim ◽  
Junaedi Fahmi ◽  
Rike Adelia ◽  
Fandy Akhmad Nur Febryanto ◽  
...  

Sign in / Sign up

Export Citation Format

Share Document