Logo image
建構一個以共時與歷時語言研究為導向的歷史語料庫
Journal article

建構一個以共時與歷時語言研究為導向的歷史語料庫

培泉 魏, 樸森 譚, 承慧 劉, 居仁 黃 and 朝奮 孫
華藝線上圖書館 中文計算語言學期刊, Vol.2(1), pp.131-145
01/02/1997

Abstract

語料庫;詞彙庫;詞類;標記;檢索;古代漢語;中古漢語;近代漢語
The Academia Sinica Ancient Chinese Corpus is designed for linguistic research. The corpus contains ancient texts that are selected because of their usefulness in grammatical and lexical studies, as well as an inspection program with keyword searching, statistics, and collocation functions. The corpus is divided into three subcorpora according to stages of grammatical developments, thus both synchronic and diachronic studies can be performed on them. Their current sizes are as follows: a. Old Chinese subcorpus (from pre-Qin to Pre-Han): 5,128,068 characters. b. Middle Chinese subcorpus (from Late Han to the Six Dynasties): 8,101,662 characters. c. Early Mandarin Chinese subcorpus (from Tang to Ching): 4,406,381 characters. A great portion of the texts from the Old Chinese subcorpus (4,497,051 characters) has been textually classified and marked-up according to their source books, author, text genre etc. A substantive part (520,794 characters) of the same subcorpus has also been segmented into words, which are in turn given part-of-speech tagging. results of the above two tasks form the basis of our Old Chinese Lexical Database.

Metrics

1 Record Views

Details

Logo image