摘要
•Discover consistent and inconsistent cross-modal feature transitions with multi-modal FAS datasets.•Propose cross-modal feature transition learning for multi-modal FAS.•Design complementary learning to infer IR- and depth-like features from RGB.•Achieve state-of-the-art performance across multiple FAS protocols.
Multi-modal face anti-spoofing (FAS) aims to detect genuine human presence by extracting discriminative liveness cues from multiple modalities, such as RGB, infrared (IR), and depth images, to enhance the robustness of biometric authentication systems. However, because data from different modalities are typically captured by various camera sensors and under diverse environmental conditions, multi-modal FAS often exhibits significantly larger distribution discrepancies across training and testing domains compared to single-modal FAS. Furthermore, during the inference stage, multi-modal FAS confronts even greater challenges when one or more modalities are unavailable or inaccessible. To address these issues, we propose a Cross-modal Transition-guided Network (CTNet) for robust multi-modal FAS. Our motivation stems from that, within a single modality, live faces exhibit smaller visual variations than spoof faces, and cross-modal feature transitions are more consistent for live samples than for spoof ones. Upon this insight, we propose learning consistent cross-modal feature transitions among live samples to construct a generalized feature space. Next, we introduce learning inconsistent cross-modal transitions between live and spoof samples to effectively detect out-of-distribution (OOD) attacks during inference. To further address the issue of missing modalities, we propose learning complementary IR and depth features from the RGB modality as auxiliary modalities. Extensive experiments demonstrate that the proposed CTNet outperforms previous multi-modal FAS methods across most protocols.