Logo image
Toward Automatic Generation of Transcript from Spoken Lectures: The 'Dream of the Red Chamber' Series
Conference paper

Toward Automatic Generation of Transcript from Spoken Lectures: The 'Dream of the Red Chamber' Series

Tzu-Han Lin, Kuan-Lin Lee, Hsin-Yun Chung, Fu-Hai Frank Wu, Jui-Chu Li, Tung-Lung Li, Shih-Lung Lo and Yi-Wen Liu
2022 25th Conference of the Oriental COCOSDA International Committee for the Co-Ordination and Standardisation of Speech Databases and Assessment Techniques, O-COCOSDA 2022 - Proceedings
2022

Abstract

lecture corpus speech recognition text simplification transfer learning Computer Science Applications Computer Vision and Pattern Recognition Information Systems Linguistics and Language Information Systems and Management Artificial Intelligence
We present a corpus of Mandarin literature lectures. The recording device, environment, and the topics covered in the lectures are described briefly; then, we developed speech recognition and text simplification approaches for automatic transcription of the lectures. Given that the size of this corpus is relatively small, we applied transfer learning on AISHELL-1 by varying the number of transferred layers and fine-Tuning the learning rates to improve the accuracy of speech recognition. Experimental results showed that a character error rate (CER) of 15.83% could be obtained. Additionally, by reducing the perplexity of the language model and the number of out-of-vocabulary words, the CER improved by another 0.29%. Further, to improve the fluency of transcription, we chose to deal with punctuation and pleonasm. Accuracy of 81% and 91.5% were reported by professional judges on adding punctuation and deleting pleonasm, respectively.

Metrics

1 Record Views

Details

Logo image