摘要
Affective speech technology aspires to equip machines with the ability to sense, interpret, and generate emotionally expressive speech, enabling empathetic assistants, social robots, and digital health companions. Large Audio/Speech Language Models (LALMs/SpeechLMs) now dominate this space: a single model can perform speech recognition, affect detection, and emotion-controlled synthesis, achieving impressive zero-shot generalization. However, we argue that LALMs are not yet internationalized: culturally grounded affect is misread when training data are skewed, leading to mis-recognition of affect, culturally inappropriate responses, and uneven user experiences. This paper surveys the current state of affective speech processing with LALMs, cataloging leading models, their sensing-to-synthesis capabilities, and the databases and metrics used for evaluation. We identify the key obstacle to responsible deployment: the heterogeneity of human vocal expression across cultures, which manifests as data scarcity, model bias, and evaluation blind spots. To address this gap, we propose a research agenda comprising: (i) systematic analysis of cultural variation in vocal affect, (ii) computational strategies for contextualizing LALMs toward culturally sensitive emotion processing, and (iii) benchmarks featuring balanced corpora and culture-aware metrics. By charting these directions, we aim to advance affective speech technology that is globally robust, socially responsible, and truly inclusive. The overall concept is depicted in Figure 1.