摘要
Modeling interaction is a growing priority across language, speech, and behavior-based AI. Beyond isolated utterances , interaction involves evolving turn dynamics, speaker-specific context, and tightly coupled multimodal and multilingual cues. However, existing datasets rarely support integrated modeling of these factors within natural conversations. As a result, many studies remain limited in their ability to reflect culturally grounded and socially adaptive behaviors. In this paper, we systematically examine representative multimodal and multilingual interaction datasets for key structural gaps that hinder progress in real-world interaction modeling. We highlight three essential but under-addressed dimensions in current dataset design: transferability, diversity and granularity. These limitations reduce the capacity of datasets to support generalization across users, represent a broad range of interaction styles, or capture fine-grained social dynamics. We advocate a shift in dataset development priorities: from convenience-driven collection to interaction-centered design. This research bridges the disconnect between current data practices and the emerging needs of interactive AI.