Voice Conversion Using Spectral Mapping and TD-PSOLA

Abstract The Gaussian mixture model (GMM) method is popular and efficient for voice conversion (VC), but it is often subject to overfitting. In this paper, the principal component regression (PCR) method is adopted for the spectral mapping between source speech and target speech, and the numbers of principal components are adjusted properly to prevent the overfitting. Then, in order to better model the nonlinear relationships between the source speech and target speech, the kernel principal component regression (KPCR) method is also proposed. Moreover, a KPCR combined with GMM method is further proposed to improve the accuracy of conversion. In addition, the discontinuity and oversmoothing problems of the traditional GMM method are also addressed. On the one hand, in order to solve the discontinuity problem, the adaptive median filter is adopted to smooth the posterior probabilities. On the other hand, the two mixture components with higher posterior probabilities for each frame are chosen for VC to reduce the oversmoothing problem. Finally, the objective and subjective experiments are carried out, and the results demonstrate that the proposed approach shows greatly better performance than the GMM method. In the objective tests, the proposed method shows lower cepstral distances and higher identification rates than the GMM method. While in the subjective tests, the proposed method obtains higher scores of preference and perceptual quality.

Download Full-text

Spectral Mapping Using Prior Re-Estimation of i-Vectors and System Fusion for Voice Conversion

IEEE/ACM Transactions on Audio Speech and Language Processing ◽

10.1109/taslp.2017.2743620 ◽

2017 ◽

Vol 25 (11) ◽

pp. 2071-2084 ◽

Cited By ~ 2

Author(s):

Monisankha Pal ◽

Goutam Saha

Keyword(s):

Voice Conversion ◽

Spectral Mapping

Download Full-text

An evaluation of voice conversion with neural network spectral mapping models and WaveNet vocoder

APSIPA Transactions on Signal and Information Processing ◽

10.1017/atsip.2020.24 ◽

2020 ◽

Vol 9 ◽

Author(s):

Patrick Lumban Tobing ◽

Yi-Chiao Wu ◽

Tomoki Hayashi ◽

Kazuhiro Kobayashi ◽

Tomoki Toda

Keyword(s):

Neural Network ◽

Statistical Models ◽

Voice Conversion ◽

Spectral Mapping ◽

High Quality ◽

Spectral Modeling ◽

Mixture Density ◽

Waveform Generation ◽

Quality Degradation ◽

Mapping Models

This paper presents an evaluation of parallel voice conversion (VC) with neural network (NN)-based statistical models for spectral mapping and waveform generation. The NN-based architectures for spectral mapping include deep NN (DNN), deep mixture density network (DMDN), and recurrent NN (RNN) models. WaveNet (WN) vocoder is employed as a high-quality NN-based waveform generation. In VC, though, owing to the oversmoothed characteristics of estimated speech parameters, quality degradation still occurs. To address this problem, we utilize post-conversion for the converted features based on direct waveform modifferential and global variance postfilter. To preserve the consistency with the post-conversion, we further propose a spectrum differential loss for the spectral modeling. The experimental results demonstrate that: (1) the RNN-based spectral modeling achieves higher accuracy with a faster convergence rate and better generalization compared to the DNN-/DMDN-based models; (2) the RNN-based spectral modeling is also capable of producing less oversmoothed spectral trajectory; (3) the use of proposed spectrum differential loss improves the performance in the same-gender conversions; and (4) the proposed post-conversion on converted features for the WN vocoder in VC yields the best performance in both naturalness and speaker similarity compared to the conventional use of WN vocoder.

Download Full-text

Voice Conversion With CycleRNN-Based Spectral Mapping and Finely Tuned WaveNet Vocoder

IEEE Access ◽

10.1109/access.2019.2955978 ◽

2019 ◽

Vol 7 ◽

pp. 171114-171125 ◽

Cited By ~ 2

Author(s):

Patrick Lumban Tobing ◽

Yi-Chiao Wu ◽

Tomoki Hayashi ◽

Kazuhiro Kobayashi ◽

Tomoki Toda

Keyword(s):

Voice Conversion ◽

Spectral Mapping

Download Full-text

Local partial least square regression for spectral mapping in voice conversion

2013 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference ◽

10.1109/apsipa.2013.6694332 ◽

2013 ◽

Cited By ~ 1

Author(s):

Xiaohai Tian ◽

Zhizheng Wu ◽

Eng Siong Chng

Keyword(s):

Partial Least Square ◽

Least Square ◽

Partial Least Square Regression ◽

Voice Conversion ◽

Spectral Mapping ◽

Least Square Regression

Download Full-text

Noise-Robust Voice Conversion Based on Sparse Spectral Mapping Using Non-negative Matrix Factorization

IEICE Transactions on Information and Systems ◽

10.1587/transinf.e97.d.1411 ◽

2014 ◽

Vol E97.D (6) ◽

pp. 1411-1418 ◽

Cited By ~ 9

Author(s):

Ryo AIHARA ◽

Ryoichi TAKASHIMA ◽

Tetsuya TAKIGUCHI ◽

Yasuo ARIKI

Keyword(s):

Matrix Factorization ◽

Voice Conversion ◽

Spectral Mapping ◽

Noise Robust ◽

Non Negative Matrix Factorization

Download Full-text

Voice conversion based on Gaussian mixture modules with Minimum Distance Spectral Mapping

2015 5th International Conference on Information Science and Technology (ICIST) ◽

10.1109/icist.2015.7288996 ◽

2015 ◽

Cited By ~ 1

Author(s):

Gui Jin ◽

Michael T. Johnson ◽

Jia Liu ◽

Xiaokang Lin

Keyword(s):

Minimum Distance ◽

Gaussian Mixture ◽

Voice Conversion ◽

Spectral Mapping

Download Full-text

Introduction of Spectral Mapping through Transmission Grating, Derivative Technique of Photon Emission

ISTFA 2014: Conference Proceedings from the 40th International Symposium for Testing and Failure Analysis ◽

10.31399/asm.cp.istfa2014p0115 ◽

2014 ◽

Author(s):

Thierry Parrassin ◽

Sylvain Dudit ◽

Michel Vallet ◽

Antoine Reverdy ◽

Hervé Deslandes

Keyword(s):

Integrated Circuits ◽

Failure Analysis ◽

Light Emission ◽

Photon Emission ◽

Optical Path ◽

Spectral Mapping ◽

Additional Information ◽

Transmission Grating ◽

Emission System ◽

Advance Knowledge

Abstract By adding a transmission grating into the optical path of our photon emission system and after calibration, we have completed several failure analysis case studies. In some cases, additional information on the emission sites is provided, as well as understanding of the behavior of transistors that are associated to the fail site. The main application of the setup is used for finding and differentiating easily related emission spots without advance knowledge in light emission mechanisms in integrated circuits.

Download Full-text

Voice Conversion Using Spectral Mapping and TD-PSOLA

Spectral Mapping Using Artificial Neural Networks for Voice Conversion

Voice Conversion using GMM with Minimum Distance Spectral Mapping Plus Amplitude Scaling

Spectral Mapping Using Kernel Principal Components Regression for Voice Conversion

Spectral Mapping Using Prior Re-Estimation of i-Vectors and System Fusion for Voice Conversion

An evaluation of voice conversion with neural network spectral mapping models and WaveNet vocoder

Voice Conversion With CycleRNN-Based Spectral Mapping and Finely Tuned WaveNet Vocoder

Local partial least square regression for spectral mapping in voice conversion

Noise-Robust Voice Conversion Based on Sparse Spectral Mapping Using Non-negative Matrix Factorization

Voice conversion based on Gaussian mixture modules with Minimum Distance Spectral Mapping

Introduction of Spectral Mapping through Transmission Grating, Derivative Technique of Photon Emission

Export Citation Format