Abstract
This work investigates automatic speech recognition (ASR) and cross-lingual transcription for Atayal, an endangered Austronesian language spoken in Taiwan, by leveraging the Whisper-medium pre-trained model. We finetuned two task-specific systems: one for direct Atayal speech-to-text transcription and another for Atayal speech-toChinese translation. Both models were trained on a bilingual Atayal-Mandarin corpus and demonstrated strong performance, achieving character error rates (CER) of 3.617 % for Atayal transcription and 5.585 % for Mandarin output. The results show the effectiveness of fine-tuning large-scale ASR models for low-resource languages and demonstrate their potential to aid the documentation and preservation of endangered languages.