Research on Speech Synthesis Based on Mixture Alignment Mechanism.

Abstract

In recent years, deep learning-based speech synthesis has attracted a lot of attention from the machine learning and speech communities. In this paper, we propose Mixture-TTS, a non-autoregressive speech synthesis model based on mixture alignment mechanism. Mixture-TTS aims to optimize the alignment information between text sequences and mel-spectrogram. Mixture-TTS uses a linguistic encoder based on soft phoneme-level alignment and hard word-level alignment approaches, which explicitly extract word-level semantic information, and introduce pitch and energy predictors to optimally predict the rhythmic information of the audio. Specifically, Mixture-TTS introduces a post-net based on a five-layer 1D convolution network to optimize the reconfiguration capability of the mel-spectrogram. We connect the output of the decoder to the post-net through the residual network. The mel-spectrogram is converted into the final audio by the HiFi-GAN vocoder. We evaluate the performance of the Mixture-TTS on the AISHELL3 and LJSpeech datasets. Experimental results show that Mixture-TTS is somewhat better in alignment information between the text sequences and mel-spectrogram, and is able to achieve high-quality audio. The ablation studies demonstrate that the structure of Mixture-TTS is effective.

Show Full Text

Research on Speech Synthesis Based on Mixture Alignment Mechanism.

Researchers

Journal

Modalities

Models

Abstract

MCSC-Net: COVID-19 detection using deep-Q-neural network classification with RFNN-based hybrid whale optimization.

Deep learning in forensic gunshot wound interpretation-a proof-of-concept study.

A Classification Method for Thoracolumbar Vertebral Fractures due to Basketball Sports Injury Based on Deep Learning.

FLImBrush: dynamic visualization of intraoperative free-hand fiber-based fluorescence lifetime imaging.

EmotionNet Nano: An Efficient Deep Convolutional Neural Network Design for Real-Time Facial Expression Recognition.

Deep Learning Techniques for Vehicle Detection and Classification from Images/Videos: A Survey.

Leave a Reply Cancel reply

Researchers

Journal

Modalities

Models

Abstract

Similar Posts

Leave a Reply Cancel reply