Learning string edit distance

dc.creatorRistad, Eric Sven
dc.creatorYianilos, Peter N.
dc.date1996-10-29
dc.date1997-11-02
dc.date.accessioned2026-07-07T08:58:43Z
dc.date.available2026-07-07T08:58:43Z
dc.descriptionIn many applications, it is necessary to determine the similarity of two strings. A widely-used notion of string similarity is the edit distance: the minimum number of insertions, deletions, and substitutions required to transform one string into the other. In this report, we provide a stochastic model for string edit distance. Our stochastic model allows us to learn a string edit distance function from a corpus of examples. We illustrate the utility of our approach by applying it to the difficult problem of learning the pronunciation of words in conversational speech. In this application, we learn a string edit distance with one fourth the error rate of the untrained Levenshtein distance. Our approach is applicable to any string classification problem that may be solved using a similarity function against a database of labeled prototypes. Keywords: string edit distance, Levenshtein distance, stochastic transduction, syntactic pattern recognition, prototype dictionary, spelling correction, string correction, string similarity, string classification, speech recognition, pronunciation modeling, Switchboard corpus.
dc.descriptionhttp://www.cs.princeton.edu/~ristad/papers/pu-532-96.ps.gz
dc.identifierhttps://arxiv.org/abs/cmp-lg/9610005
dc.identifierhttp://arxiv.org/abs/cmp-lg/9610005
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/147435
dc.subjectComputation and Language
dc.titleLearning string edit distance
dc.typetext

Files

Collections