Abstract
<p style="text-align:justify"><span style="font-size:12pt"><span style="text-justify:inter-ideograph"><span style="font-family:Calibri,sans-serif"><b><span lang="EN-US" style="font-size:14.0pt"><span style="font-family:"Times New Roman",serif">187<sup>th</sup> Acoustical Society of America (ASA) Meeting</span></span></b></span></span></span></p><p style="text-align:justify"><span style="font-size:12pt"><span style="text-justify:inter-ideograph"><span style="font-family:Calibri,sans-serif"><b><span lang="EN-US" style="font-size:14.0pt"><span style="font-family:"Times New Roman",serif">Title </span></span></b></span></span></span><span style="font-size:12pt"><span style="text-justify:inter-ideograph"><span style="font-family:Calibri,sans-serif"><span lang="EN-US" style="font-family:"Times New Roman",serif">A Multichannel Audio Tagging and Localization System for Home Surveillance</span></span></span></span></p><p style="text-align:justify"><strong><span style="font-size:12pt"><span style="text-justify:inter-ideograph"><span style="font-family:Calibri,sans-serif"><span lang="EN-US" style="font-family:"Times New Roman",serif">Abstract</span></span></span></span></strong></p><p style="text-align:justify"><span style="font-size:12pt"><span style="text-justify:inter-ideograph"><span style="font-family:Calibri,sans-serif"><span lang="EN-US" style="font-family:"Times New Roman",serif">Audio Tagging (AT) is a critical technique in smart home applications, enabling continuous monitoring of specific sound events for subsequent surveillance systems. Although Deep Neural Networks (DNNs) can provide promising tagging results, the performance degrades significantly in adverse acoustic environments. In addition, if the AT system can provide both the tagging and the location results, it can greatly enhance the system’s ability to capture sound events. To this end, a multichannel audio tagging system based on Convolutional Neural Network (CNN) is proposed to simultaneously tag and localize the sound event. Interchannel Phase Differences (IPDs) between each pair of microphones are used as the spatial feature for the input to the model. Instead of predicting the angle or zone index, the proposed system outputs a unit vector pointing to the sound event for localization. A novel loss function is introduced that computes the square of the cosine similarity between the ground truth and the estimated unit vector, allowing for more accurate localization performance. Experimental results show that the proposed system outperforms the single-channel baseline under strong interference, making it well suited for real-world applications. In addition, the model trained with the proposed loss function can greatly improve the localization accuracy.</span></span></span></span></p>