Using compression to identify acronyms in text

dc.creatorYeates, Stuart
dc.creatorBainbridge, David
dc.creatorWitten, Ian H.
dc.date2000-07-04
dc.date.accessioned2026-07-07T03:16:20Z
dc.date.available2026-07-07T03:16:20Z
dc.descriptionText mining is about looking for patterns in natural language text, and may be defined as the process of analyzing text to extract information from it for particular purposes. In previous work, we claimed that compression is a key technology for text mining, and backed this up with a study that showed how particular kinds of lexical tokens---names, dates, locations, etc.---can be identified and located in running text, using compression models to provide the leverage necessary to distinguish different token types (Witten et al., 1999)
dc.description10 pages. A short form published in DCC2000
dc.identifierhttps://arxiv.org/abs/cs/0007003
dc.identifierhttp://arxiv.org/abs/cs/0007003
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/30312
dc.subjectDigital Libraries
dc.subjectInformation Retrieval
dc.subjectH.3.7
dc.titleUsing compression to identify acronyms in text
dc.typetext

Files

Collections