Sound Source Separation Mechanisms of Different Deep Networks Explained from the Perspective of Auditory Perception

Thanks to the development of deep learning, various sound source separation networks have been proposed and made significant progress. However, the study on the underlying separation mechanisms is still in its infancy. In this study, deep networks are explained from the perspective of auditory perception mechanisms. For separating two arbitrary sound sources from monaural recordings, three different networks with different parameters are trained and achieve excellent performances. The networks’ output can obtain an average scale-invariant signal-to-distortion ratio improvement (SI-SDRi) higher than 10 dB, comparable with the human performance to separate natural sources. More importantly, the most intuitive principle—proximity—is explored through simultaneous and sequential organization experiments. Results show that regardless of network structures and parameters, the proximity principle is learned spontaneously by all networks. If components are proximate in frequency or time, they are not easily separated by networks. Moreover, the frequency resolution at low frequencies is better than at high frequencies. These behavior characteristics of all three networks are highly consistent with those of the human auditory system, which implies that the learned proximity principle is not accidental, but the optimal strategy selected by networks and humans when facing the same task. The emergence of the auditory-like separation mechanisms provides the possibility to develop a universal system that can be adapted to all sources and scenes.

Download Full-text

An adaptive time-frequency resolution approach for Non-negative Matrix Factorization based single channel sound source separation

2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) ◽

10.1109/icassp.2011.5946388 ◽

2011 ◽

Cited By ~ 7

Author(s):

Serap Kirbiz ◽

Paris Smaragdis

Keyword(s):

Sound Source ◽

Matrix Factorization ◽

Single Channel ◽

Source Separation ◽

Frequency Resolution ◽

Time Frequency ◽

Sound Source Separation ◽

Adaptive Time ◽

Non Negative Matrix Factorization

Download Full-text

A multichannel learning-based approach for sound source separation in reverberant environments

EURASIP Journal on Audio Speech and Music Processing ◽

10.1186/s13636-021-00227-2 ◽

2021 ◽

Vol 2021 (1) ◽

Author(s):

You-Siang Chen ◽

Zi-Jie Lin ◽

Mingsian R. Bai

Keyword(s):

Frequency Domain ◽

Sound Source ◽

Signal To Noise Ratio ◽

Source Separation ◽

Objective Evaluation ◽

Linear Mapping ◽

Invariant Mean ◽

Scale Invariant ◽

Sound Source Separation ◽

Reverberant Field

AbstractIn this paper, a multichannel learning-based network is proposed for sound source separation in reverberant field. The network can be divided into two parts according to the training strategies. In the first stage, time-dilated convolutional blocks are trained to estimate the array weights for beamforming the multichannel microphone signals. Next, the output of the network is processed by a weight-and-sum operation that is reformulated to handle real-valued data in the frequency domain. In the second stage, a U-net model is concatenated to the beamforming network to serve as a non-linear mapping filter for joint separation and dereverberation. The scale invariant mean square error (SI-MSE) that is a frequency-domain modification from the scale invariant signal-to-noise ratio (SI-SNR) is used as the objective function for training. Furthermore, the combined network is also trained with the speech segments filtered by a great variety of room impulse responses. Simulations are conducted for comprehensive multisource scenarios of various subtending angles of sources and reverberation times. The proposed network is compared with several baseline approaches in terms of objective evaluation matrices. The results have demonstrated the excellent performance of the proposed network in dereverberation and separation, as compared to baseline methods.

Download Full-text