Abstract
In the knowledge-centric environment, enterprise knowledge acquisition, storage, management and reuse are the typical issues for enterprises to maintain their advantages in the global market. However, the present knowledge management techniques focus mainly on document search, version control and authorization. The contents and structure of documents that reveal the critical knowledge are rarely concerned. On the other hand, owing to the popularity of the Internet technology, more and more enterprise knowledge is exchanged and reused over Internet. In order to effectively explore the critical information in the free-from documents, a model for document structure analysis is developed in this research. In the proposed methodology, based on keyword frequency and correlation, document fragmentation and pattern analysis algorithms are utilized to analyze the document components and structure. Using the knowledge ontology defined based on the RDF syntax, the document components are then parsed into semantic structure. In addition to the document content analysis model, a prototype system is also developed and an IP management case is provided to verify the feasibility and effectiveness of the model. This research aims at developing an applicable approach to transform the free-form documents into structured semantic representation. As a result, the goal of automatic knowledge extraction and reused can be fulfilled and efficiency of enterprise knowledge management can be significantly improved.