A Maximum Entropy Approach to Identifying Sentence Boundaries

dc.creatorReynar, Jeffrey C.
dc.creatorRatnaparkhi, Adwait
dc.date1997-04-09
dc.date.accessioned2026-07-07T09:10:45Z
dc.date.available2026-07-07T09:10:45Z
dc.descriptionWe present a trainable model for identifying sentence boundaries in raw text. Given a corpus annotated with sentence boundaries, our model learns to classify each occurrence of ., ?, and ! as either a valid or invalid sentence boundary. The training procedure requires no hand-crafted rules, lexica, part-of-speech tags, or domain-specific information. The model can therefore be trained easily on any genre of English, and should be trainable on any other Roman-alphabet language. Performance is comparable to or better than the performance of similar systems, but we emphasize the simplicity of retraining for new domains.
dc.description4 pages, uses aclap.sty and covingtn.sty
dc.identifierhttps://arxiv.org/abs/cmp-lg/9704002
dc.identifierhttp://arxiv.org/abs/cmp-lg/9704002
dc.identifierProceedings of the 5th ANLP Conference, 1997
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/151443
dc.subjectComputation and Language
dc.titleA Maximum Entropy Approach to Identifying Sentence Boundaries
dc.typetext

Files

Collections