Mostly-Unsupervised Statistical Segmentation of Japanese Kanji Sequences
| dc.creator | Ando, Rie Kubota | |
| dc.creator | Lee, Lillian | |
| dc.date | 2002-05-10 | |
| dc.date.accessioned | 2026-07-07T03:18:23Z | |
| dc.date.available | 2026-07-07T03:18:23Z | |
| dc.description | Given the lack of word delimiters in written Japanese, word segmentation is generally considered a crucial first step in processing Japanese texts. Typical Japanese segmentation algorithms rely either on a lexicon and syntactic analysis or on pre-segmented data; but these are labor-intensive, and the lexico-syntactic techniques are vulnerable to the unknown word problem. In contrast, we introduce a novel, more robust statistical method utilizing unsegmented training data. Despite its simplicity, the algorithm yields performance on long kanji sequences comparable to and sometimes surpassing that of state-of-the-art morphological analyzers over a variety of error metrics. The algorithm also outperforms another mostly-unsupervised statistical algorithm previously proposed for Chinese. Additionally, we present a two-level annotation scheme for Japanese to incorporate multiple segmentation granularities, and introduce two novel evaluation metrics, both based on the notion of a compatible bracket, that can account for multiple granularities simultaneously. | |
| dc.description | 22 pages. To appear in Natural Language Engineering | |
| dc.identifier | https://arxiv.org/abs/cs/0205009 | |
| dc.identifier | http://arxiv.org/abs/cs/0205009 | |
| dc.identifier | Natural Language Engineering 9 (2), pp. 127--149, 2003 | |
| dc.identifier | doi:10.1017/S1351324902002954 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/31091 | |
| dc.subject | Computation and Language | |
| dc.subject | I.2.7 | |
| dc.title | Mostly-Unsupervised Statistical Segmentation of Japanese Kanji Sequences | |
| dc.type | text |