摘要
•Introduces three new strategies to enhance LDCformer in FAS.•Dual-attention supervision enables fine-grained liveness learning.•Self-challenging supervision improves partial spoof detection.•Transitional triplet mining boosts cross-domain generalization.•Achieves SOTA results on multiple FAS benchmarks.
Face anti-spoofing (FAS) heavily relies on identifying live/spoof discriminative features to counter face presentation attacks. Recently, we proposed LDCformer to successfully incorporate the Learnable Descriptive Convolution (LDC) into ViT to model long-range dependency of locally descriptive features for FAS. In this paper, we propose three novel training strategies to effectively enhance the training of LDCformer to largely boost its feature characterization capability. The first strategy, dual-attention supervision, is developed to learn fine-grained liveness features guided by regional live/spoof attentions. The second strategy, self-challenging supervision, is designed to enhance feature discriminability by generating challenging training samples. In addition, we propose a third training strategy, transitional triplet mining strategy, through narrowing the cross-domain gap while preserving the transitional relationship between live and spoof features, to enlarge the domain-generalization capability of LDCformer. Extensive experiments show that LDCformer, under joint supervision of these three novel training strategies, significantly outperforms previous methods and verifies the effectiveness and potential applicability of these strategies to other network architectures.