Abstract
Emotion recognition has been successfully applied in many fields. It is believed that features extracted from each timing-level can provide different information of the emotional speech signals and therefore can compensate one another. In order to achieve a promising recognition accuracy, several methods for combining features extracted from different timing-levels are proposed in this thesis, including likelihood combination, weighted likelihood combination, raw feature combination and partial raw feature combination. We extracted spectrum features and prosodic features for frame-level features, and low-level descriptors (LLDs) for segment-level features and utterance-level features. The Berlin Emotion Database and eNTERFACE emotional database are used in the experiments. Compared with conventional one or two timing-level features, the combination of three timing-level features shows higher recognition rate.