Logo image
Learning-based robust speaker counting and separation with the aid of spatial coherence
期刊文章   開放取用(OA)   同儕審查

Learning-based robust speaker counting and separation with the aid of spatial coherence

Yicheng HsuMingsian R. Bai
Eurasip Journal on Audio, Speech, and Music Processing, 卷.2023(1), 36
12/2023

摘要

Multichannel blind source separation Neural network Spatial coherence Speaker counting and separation Acoustics and Ultrasonics Electrical and Electronic Engineering
A three-stage approach is proposed for speaker counting and speech separation in noisy and reverberant environments. In the spatial feature extraction, a spatial coherence matrix (SCM) is computed using whitened relative transfer functions (wRTFs) across time frames. The global activity functions of each speaker are estimated from a simplex constructed using the eigenvectors of the SCM, while the local coherence functions are computed from the coherence between the wRTFs of a time-frequency bin and the global activity function-weighted RTF of the target speaker. In speaker counting, we use the eigenvalues of the SCM and the maximum similarity of the interframe global activity distributions between two speakers as the input features to the speaker counting network (SCnet). In speaker separation, a global and local activity-driven network (GLADnet) is used to extract each independent speaker signal, which is particularly useful for highly overlapping speech signals. Experimental results obtained from the real meeting recordings show that the proposed system achieves superior speaker counting and speaker separation performance compared to previous publications without the prior knowledge of the array configurations.

檔案與連結 (1)

url
https://doi.org/10.1186/s13636-023-00298-3檢視
已出版(紀錄版本) 開放

相關連結

指標

1 檢視次數

詳細資料

Logo image