Abstract
The number of children diagnosed with Autism Spectrum Disorder (ASD) is continually increasing. However, diagnosing autism is not straightforward; the process is lengthy and complex. Previous research has indicated that children with autism exhibit deficits in using language within social contexts, such as storytelling skills. Additionally, children with autism may display distinct patterns in certain acoustic features compared to typically developing (TD) children. With the advancement of computational models, we aim to employ deep neural network models to rapidly analyze the acoustic features of children’s narratives for detecting autism. In this study, we collected narrative data from 12 children with autism and 19 TD children using Module 3 of the standardized tool ADOS-2 (Autism Diagnostic Observation Schedule, Second Edition). We then represented the acoustic features using Mel-Frequency Cepstral Coefficients (MFCCs) and employed computational models for training and classification. Moreover, we identified 10 low-level descriptors (LLDs) in our dataset through t-tests that showed significant differences between ASD and TD children. We combined these 10 LLDs with MFCCs as inputs to the model and achieved an F1 score of 89.4%. Upon further analysis of these 10 LLDs, we found that they indeed represent the speech characteristic differences between ASD and TD children. This further provides interpretability of our model in identifying autistic tendencies based on the speech characteristic differences.