Currently, there are few available tools to separate ancient Japanese sentences into words. Therefore, it is difficult to extract archaic Japanese words from Japanese ancient writings. We propose a method of word segmentation for Japanese ancient writings. We calculate the likelihood of character n-grams to be words, and extract character n-grams with higher likelihood as archaic Japanese words. We conducted word separation experiments using the term likelihood with the proposed method.
展开▼