Logo image
Audience-Aware Co-speech Gesture Generation in Public Speaking via Anticipation Tokens
會議論文

Audience-Aware Co-speech Gesture Generation in Public Speaking via Anticipation Tokens

Huan Yu Chen, Woan-Shiuan Chien 和 Chi-Chun Lee
ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (Barcelona, Spain, 03/05/2026–08/05/2026)
03/05/2026

摘要

Co-speech gestures are vital in both social interaction and human–computer communication, carrying semantic, affective, and regulatory signals. Prior work largely frames public speaking as a single-speaker task, but this view neglects the audience, whose presence makes performance possible. We argue that public speaking involves an asymmetrical yet interactive dynamic, where the speaker actively anticipates and elicits audience responses. To capture this dynamic, we introduce a gesture generation framework that integrates audience anticipation via pre-response tokens representing cues such as forthcoming laughter. Experiments show that jointly modeling speaker delivery and audience anticipation improves the realism of generated gestures. Furthermore, our ablation study reveals that conditioning anticipation on speech representations yields more stable benefits than injecting it directly into gesture diffusion. These findings highlight the importance of audience-aware modeling for advancing gesture synthesis in public speaking scenarios.

相關連結

指標

1 檢視次數

詳細資料

Logo image