Indexing Schemes for Similarity Search In Datasets of Short Protein Fragments

dc.creatorStojmirovic, Aleksandar
dc.creatorPestov, Vladimir
dc.date2003-09-05
dc.date2007-02-09
dc.date.accessioned2026-07-07T08:27:19Z
dc.date.available2026-07-07T08:27:19Z
dc.descriptionWe propose a family of very efficient hierarchical indexing schemes for ungapped, score matrix-based similarity search in large datasets of short (4-12 amino acid) protein fragments. This type of similarity search has importance in both providing a building block to more complex algorithms and for possible use in direct biological investigations where datasets are of the order of 60 million objects. Our scheme is based on the internal geometry of the amino acid alphabet and performs exceptionally well, for example outputting 100 nearest neighbours to any possible fragment of length 10 after scanning on average less than one per cent of the entire dataset.
dc.description34 pages, 12 figures, 4 tables - Timings for experiments added upon referees' request, and a number of less substantial modifications made
dc.identifierhttps://arxiv.org/abs/cs/0309005
dc.identifierhttp://arxiv.org/abs/cs/0309005
dc.identifierInformation Systems 32 (2007), 1145-1165
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/137231
dc.subjectData Structures and Algorithms
dc.subjectBiomolecules
dc.subjectH.3.1; J.3
dc.titleIndexing Schemes for Similarity Search In Datasets of Short Protein Fragments
dc.typetext

Files

Collections