Logo image
Cross-Modal Semantic Alignment Via Ensemble Audio-Text Features for XACLE Challenge
會議論文

Cross-Modal Semantic Alignment Via Ensemble Audio-Text Features for XACLE Challenge

Snehit B. Chunarkar, Krishnagiri HamzaChi-Chun Lee
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (Barcelona, Spain, 03/05/2026–08/05/2026)
03/05/2026

摘要

This paper presents an ensemble framework for predicting semantic audio-text alignment for GC-12: x-to-audio alignment (XACLE) in the ICASSP 2026: SP Grand Challenge. We leverage ensemble sets comprising carefully chosen complementary model features: M2D-CLAP, MS-CLAP, MGA-CLAP, LAION-CLAP, Whisper and DeBERTaV3; And Augmented with proximity features: cosine similarity, cosine angle, L1, and L2 norms. These diverse representations are combined and fed into optimized regressors to robustly estimate human-perceived semantic correlation scores for audio-text pairs. With the proposed approach, our submission achieves 2nd rank in the official leaderboard, highlighting significant boosts from incorporating both deep model and handcrafted proximity features with an optimized weighted two regression model approach.

相關連結

指標

1 檢視次數

詳細資料

Logo image