Robust Speaker Localization Guided by Deep Learning-Based Time-Frequency Masking
Author(s) -
Zhong-Qiu Wang,
Xueliang Zhang,
DeLiang Wang
Publication year - 2018
Publication title -
ieee/acm transactions on audio speech and language processing
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.916
H-Index - 56
eISSN - 2329-9304
pISSN - 2329-9290
DOI - 10.1109/taslp.2018.2876169
Subject(s) - monaural , computer science , robustness (evolution) , speech recognition , reverberation , direction of arrival , masking (illustration) , speech enhancement , artificial neural network , pattern recognition (psychology) , artificial intelligence , acoustics , noise reduction , telecommunications , antenna (radio) , physics , art , visual arts , biochemistry , chemistry , gene
Deep learning-based time-frequency (T-F) masking has dramatically advanced monaural (single-channel) speech separation and enhancement. This study investigates its potential for direction of arrival (DOA) estimation in noisy and reverberant environments. We explore ways of combining T-F masking and conventional localization algorithms, such as generalized cross correlation with phase transform, as...
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom