Logo image
Channel difference transformer for face anti-spoofing
期刊文章   同儕審查

Channel difference transformer for face anti-spoofing

Pei-Kai Huang, Jun-Xiong Chong, Ming-Tsung Hsu, Fang-Yu Hsu 和 Chiou-Ting Hsu
Information sciences, 卷.702, 頁.121904
06/2025
Web of Science ID: WOS:001421421500001

摘要

Domain generalization Face anti-spoofing Learning complementary channel information Vision transformer
Face anti-spoofing (FAS) aims to detect and counter facial presentation attacks by distinguishing subtle differences between live and spoof faces. Since the vision transformer (ViT) has demonstrated significant performance gains in various computer vision tasks by effectively capturing long-range dependencies, recent FAS methods have investigated the potential of leveraging the channel-wise features derived from self-attention (SA) in ViT. While channel-wise features effectively capture discriminative local attentions, the complementary information existing between different channels remained undiscovered in prior methods. In this paper, we investigate the unexplored characteristics of complementary channel information within ViT and propose to incorporate both channel-wise and complementary channel information to learn long-range and discriminative features for FAS. We design two modules, including Channel Difference Self-Attention (CDSA) and Multi-head Channel Difference Self-Attention (MCDSA), to facilitate learning complementary channel characteristics and enhancing both feature discriminability and representational capacity. Building upon CDSA and MCDSA, we propose a novel and efficient Channel Difference Transformer (CDformer) without introducing any additional parameters or increasing computation complexity of ViT. Extensive experiments conducted on five FAS benchmark datasets demonstrate that our proposed CDformer achieves state-of-the-art performance on both intra-domain and cross-domain testing scenarios.

相關連結

指標

1 檢視次數

詳細資料

Logo image