Deep learning based speaker separation and dereverberation can generalize across different languages to improve intelligibility.

Abstract

The practical efficacy of deep learning based speaker separation and/or dereverberation hinges on its ability to generalize to conditions not employed during neural network training. The current study was designed to assess the ability to generalize across extremely different training versus test environments. Training and testing were performed using different languages having no known common ancestry and correspondingly large linguistic differences-English for training and Mandarin for testing. Additional generalizations included untrained speech corpus/recording channel, target-to-interferer energy ratios, reverberation room impulse responses, and test talkers. A deep computational auditory scene analysis algorithm, employing complex time-frequency masking to estimate both magnitude and phase, was used to segregate two concurrent talkers and simultaneously remove large amounts of room reverberation to increase the intelligibility of a target talker. Significant intelligibility improvements were observed for the normal-hearing listeners in every condition. Benefit averaged 43.5% points across conditions and was comparable to that obtained when training and testing were performed both in English. Benefit is projected to be considerably larger for individuals with hearing impairment. It is concluded that a properly designed and trained deep speaker separation/dereverberation network can be capable of generalization across vastly different acoustic environments that include different languages.

Show Full Text

Deep learning based speaker separation and dereverberation can generalize across different languages to improve intelligibility.

Researchers

Journal

Modalities

Models

Abstract

A comprehensive segmentation of chest X-ray improves deep learning-based WHO radiologically confirmed pneumonia diagnosis in children.

Comparisons of deep learning and machine learning while using text mining methods to identify suicide attempts of patients with mood disorders.

An efficient learning based approach for automatic record deduplication with benchmark datasets.

FABEL: Forecasting Animal Behavioral Events with Deep Learning-Based Computer Vision.

Framework for Vehicle Make and Model Recognition-A New Large-Scale Dataset and an Efficient Two-Branch-Two-Stage Deep Learning Architecture.

End-to-end deep learning nonrigid motion-corrected reconstruction for highly accelerated free-breathing coronary MRA.

Leave a Reply Cancel reply

Researchers

Journal

Modalities

Models

Abstract

Similar Posts

Leave a Reply Cancel reply