Abstract
Intellectual property (IP) is a power tool for economic growth of country, it is also the competitive advantage of innovation for businesses. As the view of law, patent is used to protect IP sufficiently. With the growing of patent documents and different writing styles of claims in patents, patent analysis works including patent retrivel, synonym identifying and domain thesaurus building are extremely manual works. In this thesis, an approach of text classification we propose is relaxation labeling. The technique is used to classify the terminologies in patent documents. We have pre-classified the taxonomy of CMP domain from training data in advance. The terminologies and the information about relation, attribute and material have been extracted by NLP technique and regular expression. The probability and compatibility coefficients which are parameters in the relaxation labeling model have been estimated from training data. In the progress of relaxation labeling, the probability of each class for each unclassified term was updating. The most appropriate class will be obvious when the model is converged. Based on the extracted semantic information, the experiment results clearly show that relaxation labeling is sufficient for terminology classification and achieve certain accuracy. We believe that relaxation labeling might have been usage in patent analysis work.