Abstract
Owing to the booming growth of information technology, the number of digital documents has significantly increased over the Internet and within organizations. In order to enhance the performance for enterprises to manage their digital documents and domain knowledge, automatic document classification has become a key issue for enterprise knowledge management. Concerning complexity of different types of digital documents, this paper utilizes the principal component analysis (PCA) to develop an algorithm for automatic document classification. Based on PCA, representative keywords of distinct document categories can be obtained. Furthermore, according to the frequencies of representative keywords in the target document, the category of the target document can be determined. In addition to the document classification algorithm, a Web-based document classification system is also developed and a demonstration case is applied to verify the performance of the proposed approach. The attempt of this research is to enhance the accuracy and efficiency of enterprise document classification technology and to enable a self-service knowledge management mechanism in organizations.