Similarity-Based Estimation of Word Cooccurrence Probabilities

dc.creatorDagan, Ido
dc.creatorPereira, Fernando
dc.creatorLee, Lillian
dc.date1994-05-02
dc.date.accessioned2026-07-07T09:08:14Z
dc.date.available2026-07-07T09:08:14Z
dc.descriptionIn many applications of natural language processing it is necessary to determine the likelihood of a given word combination. For example, a speech recognizer may need to determine which of the two word combinations ``eat a peach'' and ``eat a beach'' is more likely. Statistical NLP methods determine the likelihood of a word combination according to its frequency in a training corpus. However, the nature of language is such that many word combinations are infrequent and do not occur in a given corpus. In this work we propose a method for estimating the probability of such previously unseen word combinations using available information on ``most similar'' words. We describe a probabilistic word association model based on distributional word similarity, and apply it to improving probability estimates for unseen word bigrams in a variant of Katz's back-off model. The similarity-based method yields a 20% perplexity improvement in the prediction of unseen bigrams and statistically significant reductions in speech-recognition error.
dc.description13 pages, to appear in proceedings of ACL-94
dc.identifierhttps://arxiv.org/abs/cmp-lg/9405001
dc.identifierhttp://arxiv.org/abs/cmp-lg/9405001
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/150649
dc.subjectComputation and Language
dc.titleSimilarity-Based Estimation of Word Cooccurrence Probabilities
dc.typetext

Files

Collections