摘要
•Propose UCros-rPPGNet, the first unsupervised cross-modal rPPG framework.•Noise-resistant learning and uncertainty fusion improve the stability of features.•Cross-spectral adapter enables missing-modality inference.•Introduce the DG-CMrPPG benchmark for evaluating unsupervised cross-modal rPPG.
Unsupervised remote photoplethysmography (rPPG) estimation aims to extract physiological signals from facial videos without relying on ground-truth labels. While single-modal unsupervised methods typically rely on the RGB modality, their performance often degrades under dynamic illumination and motion. In this paper, we propose UCros-rPPGNet, a completely unsupervised cross-modal framework that collaboratively leverages RGB and near-infrared (NIR) modalities. To ensure robust estimation, we introduce three integrated strategies: (1) Noise-resistant learning, which employs a noise learner to decouple physiological features from environmental artifacts; (2) Reliability-aware learnable fusion, which utilizes an uncertainty estimator to dynamically weight modality contributions based on signal quality; and (3) Cross-spectral translation, which employs a bi-directional adapter to maintain operational continuity during inference even when a modality is missing. Furthermore, we establish DG-CMrPPG, a new benchmark protocol specifically designed to evaluate cross-modal rPPG under diverse domain shifts and disturbances. Extensive experiments demonstrate that UCross-rPPGNet not only outperforms existing unsupervised methods but also surpasses supervised methods under several challenging scenarios, demonstrating the effectiveness of cross-modal translation and fusion for label-free physiological sensing.