Abstract
This study aims to improve the accuracy of forced alignment for Mandarin Chinese. The performance of automatic speech assessment relies on the quality of acoustic models. The first step of traditional automatic speech assessment is to perform model-based forced alignment on input recording and then compare with the ground truth of acoustic model. However, forced alignment is not accurate enough for co-articulations. Here, we focus on those co-articulations without short pauses between two syllables. For example, 蘇武(“s? w?”), 一意(“yi yi”), 無謂(“wu wei”), and so on; the syllable boundaries between the two co-articulated syllables are heavily misaligned and hence impact the quality of the assessment.We therefore propose a new approach using the characteristic of tones in Mandarin Chinese. Additional pitch features are considered to improve the accuracy of forced alignment. Three metrics are evaluated: sentence recognition rate, model ranking ratio, one-pass and two-pass alignment. The first and the second metrics are focus on model reliability. And the third metric emphasize the accuracy of alignment. The results show that the accuracy of alignment is improved while model ranking ratio is slightly down.