Abstract
In deepfake detection, diverse compression methods employed by social media platforms pose significant challenges due to varying compression rates. These variations hinder the generalization of deepfake detectors across different compression rates, termed cross-compression-rate (CCR) scenario. While existing models demonstrate robustness in cross-dataset evaluation, they often overlook the CCR scenario, which is crucial for ensuring broader applicability in real-world applications. Therefore, we introduce a novel Contrastive Physio-inspired Multi-modalities with Language guidance (CPML) framework for robust CCR deepfake detection. Our approach co-maps remote photoplethysmography (rPPG) signals and facial landmark dynamics into a common latent feature space and then aligns with a set of class prompt-guided in language semantics (e.g., real and fake classes). Specifically, we propose the Cross-Quality Similarity Learning (CQSL) strategy to learn the similarities in the rPPG signals under the variations of visual qualities. Moreover, we utilize a pre-trained vision-language model as our text encoder and propose the Cross-Modality Consistency Learning (CMCL) to pair-wisely align the multi-modal features with the textual features of the corresponding class prompts. Our extensive experiments demonstrate that the proposed achieves superior performance on both seen and unseen manipulation types and datasets, and provide a benchmark for CCR scenarios.