Similarity-Based Approaches to Natural Language Processing
| dc.creator | Lee, Lillian | |
| dc.date | 1997-08-19 | |
| dc.date.accessioned | 2026-07-07T09:10:57Z | |
| dc.date.available | 2026-07-07T09:10:57Z | |
| dc.description | This thesis presents two similarity-based approaches to sparse data problems. The first approach is to build soft, hierarchical clusters: soft, because each event belongs to each cluster with some probability; hierarchical, because cluster centroids are iteratively split to model finer distinctions. Our second approach is a nearest-neighbor approach: instead of calculating a centroid for each class, as in the hierarchical clustering approach, we in essence build a cluster around each word. We compare several such nearest-neighbor approaches on a word sense disambiguation task and find that as a whole, their performance is far superior to that of standard methods. In another set of experiments, we show that using estimation techniques based on the nearest-neighbor model enables us to achieve perplexity reductions of more than 20 percent over standard techniques in the prediction of low-frequency events, and statistically significant speech recognition error-rate reduction. | |
| dc.description | 71 pages (single-spaced) | |
| dc.identifier | https://arxiv.org/abs/cmp-lg/9708011 | |
| dc.identifier | http://arxiv.org/abs/cmp-lg/9708011 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/151516 | |
| dc.subject | Computation and Language | |
| dc.title | Similarity-Based Approaches to Natural Language Processing | |
| dc.type | text |