Importance Sampling of Word Patterns in DNA and Protein Sequences

dc.creatorChan, Hock Peng
dc.creatorZhang, Nancy R.
dc.creatorChen, Louis H. Y.
dc.date2008-11-26
dc.date.accessioned2026-07-07T12:06:08Z
dc.date.available2026-07-07T12:06:08Z
dc.descriptionMonte Carlo methods can provide accurate p-value estimates of word counting test statistics and are easy to implement. They are especially attractive when an asymptotic theory is absent or when either the search sequence or the word pattern is too short for the application of asymptotic formulae. Naive direct Monte Carlo is undesirable for the estimation of small probabilities because the associated rare events of interest are seldom generated. We propose instead efficient importance sampling algorithms that use controlled insertion of the desired word patterns on randomly generated sequences. The implementation is illustrated on word patterns of biological interest: Palindromes and inverted repeats, patterns arising from position specific weight matrices and co-occurrences of pairs of motifs.
dc.identifierhttps://arxiv.org/abs/0811.4447
dc.identifierhttp://arxiv.org/abs/0811.4447
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/208568
dc.subjectApplications
dc.subjectQuantitative Methods
dc.subjectComputation
dc.titleImportance Sampling of Word Patterns in DNA and Protein Sequences
dc.typetext

Files

Collections