A Model of Lexical Attraction and Repulsion
| dc.creator | Beeferman, Doug | |
| dc.creator | Berger, Adam | |
| dc.creator | Lafferty, John | |
| dc.date | 1997-06-13 | |
| dc.date | 1997-06-16 | |
| dc.date.accessioned | 2026-07-07T09:02:23Z | |
| dc.date.available | 2026-07-07T09:02:23Z | |
| dc.description | This paper introduces new methods based on exponential families for modeling the correlations between words in text and speech. While previous work assumed the effects of word co-occurrence statistics to be constant over a window of several hundred words, we show that their influence is nonstationary on a much smaller time scale. Empirical data drawn from English and Japanese text, as well as conversational speech, reveals that the ``attraction'' between words decays exponentially, while stylistic and syntactic contraints create a ``repulsion'' between words that discourages close co-occurrence. We show that these characteristics are well described by simple mixture models based on two-stage exponential distributions which can be trained using the EM algorithm. The resulting distance distributions can then be incorporated as penalizing features in an exponential language model. | |
| dc.description | 8 pages, LaTeX source and postscript figures for ACL/EACL'97 paper | |
| dc.identifier | https://arxiv.org/abs/cmp-lg/9706018 | |
| dc.identifier | http://arxiv.org/abs/cmp-lg/9706018 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/148611 | |
| dc.subject | Computation and Language | |
| dc.title | A Model of Lexical Attraction and Repulsion | |
| dc.type | text |