Estimating Lexical Priors for Low-Frequency Syncretic Forms

dc.creatorBaayen, Harald
dc.creatorSproat, Richard
dc.date1995-04-24
dc.date.accessioned2026-07-07T09:09:46Z
dc.date.available2026-07-07T09:09:46Z
dc.descriptionGiven a previously unseen form that is morphologically n-ways ambiguous, what is the best estimator for the lexical prior probabilities for the various functions of the form? We argue that the best estimator is provided by computing the relative frequencies of the various functions among the hapax legomena --- the forms that occur exactly once in a corpus. This result has important implications for the development of stochastic morphological taggers, especially when some initial hand-tagging of a corpus is required: For predicting lexical priors for very low-frequency morphologically ambiguous types (most of which would not occur in any given corpus) one should concentrate on tagging a good representative sample of the hapax legomena, rather than extensively tagging words of all frequency ranges.
dc.descriptionSubmitted to Computational Linguistics
dc.identifierhttps://arxiv.org/abs/cmp-lg/9504015
dc.identifierhttp://arxiv.org/abs/cmp-lg/9504015
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/151152
dc.subjectComputation and Language
dc.titleEstimating Lexical Priors for Low-Frequency Syncretic Forms
dc.typetext

Files

Collections