An Efficient Inductive Unsupervised Semantic Tagger
| dc.creator | Lua, K T | |
| dc.date | 1996-06-11 | |
| dc.date.accessioned | 2026-07-07T09:10:23Z | |
| dc.date.available | 2026-07-07T09:10:23Z | |
| dc.description | We report our development of a simple but fast and efficient inductive unsupervised semantic tagger for Chinese words. A POS hand-tagged corpus of 348,000 words is used. The corpus is being tagged in two steps. First, possible semantic tags are selected from a semantic dictionary(Tong Yi Ci Ci Lin), the POS and the conditional probability of semantic from POS, i.e., P(S|P). The final semantic tag is then assigned by considering the semantic tags before and after the current word and the semantic-word conditional probability P(S|W) derived from the first step. Semantic bigram probabilities P(S|S) are used in the second step. Final manual checking shows that this simple but efficient algorithm has a hit rate of 91%. The tagger tags 142 words per second, using a 120 MHz Pentium running FOXPRO. It runs about 2.3 times faster than a Viterbi tagger. | |
| dc.description | uuencoded postscript file. email: cmp-lg/9606012 | |
| dc.identifier | https://arxiv.org/abs/cmp-lg/9606012 | |
| dc.identifier | http://arxiv.org/abs/cmp-lg/9606012 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/151334 | |
| dc.subject | Computation and Language | |
| dc.title | An Efficient Inductive Unsupervised Semantic Tagger | |
| dc.type | text |