摘要
. Electroencephalography (EEG) and electromyography (EMG) are widely used for decoding motor intentions, yet unimodal approaches often suffer from low robustness and limited representational information. EEG-EMG hybrid brain-computer interfaces (BCIs) can bridge cortical intention and muscular execution, but challenges remain in signal alignment, fusion modeling, and clinical generalization.
. To address these issues, we propose spatio-temporal cross-attention fusion (STCAFusion), a spatiotemporal cross-attention (CA) framework that integrates EEG and EMG through multi-band based dual-branch convolutional encoders and parallel temporal and spatial CA modules. This design enables detailed modeling of inter-modal correlations across both time and space. We evaluate STCAFusion on a newly collected dataset of synchronous EEG-EMG recordings from 12 subjects, where the data were acquired under two paradigms (Reaching and Lifting) designed from daily functional upper-limb activities to emphasize directional and strength control.
. With leave-one-run-out cross-validation, STCAFusion achieves average accuracies of 84.15% and 95.22% in the two paradigms, outperforming the strongest competing EEG-EMG fusion baselines by 3.4% in the Reaching paradigm and 1.8% in the Lifting paradigm. Visualization of learned attention weights further reveals meaningful spatiotemporal EEG-EMG coupling patterns, offering insights into neural-muscular coordination patterns relevant to rehabilitation-oriented BCI design.
. These results highlight the potential of CA-based multimodal physiological signal fusion in building reliable hybrid BCI and wearable devices for upper-limb control and rehabilitation.