Corpus structure, language models, and ad hoc information retrieval

dc.creatorKurland, Oren
dc.creatorLee, Lillian
dc.date2004-05-12
dc.date.accessioned2026-07-07T03:21:15Z
dc.date.available2026-07-07T03:21:15Z
dc.descriptionMost previous work on the recently developed language-modeling approach to information retrieval focuses on document-specific characteristics, and therefore does not take into account the structure of the surrounding corpus. We propose a novel algorithmic framework in which information provided by document-based language models is enhanced by the incorporation of information drawn from clusters of similar documents. Using this framework, we develop a suite of new algorithms. Even the simplest typically outperforms the standard language-modeling approach in precision and recall, and our new interpolation algorithm posts statistically significant improvements for both metrics over all three corpora tested.
dc.descriptionTo appear, SIGIR 2004
dc.identifierhttps://arxiv.org/abs/cs/0405044
dc.identifierhttp://arxiv.org/abs/cs/0405044
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/32125
dc.subjectInformation Retrieval
dc.subjectComputation and Language
dc.subjectH.3.3; I.2.7
dc.titleCorpus structure, language models, and ad hoc information retrieval
dc.typetext

Files

Collections