PCA and K-Means decipher genome
| dc.creator | Gorban, A. N. | |
| dc.creator | Zinovyev, A. Y. | |
| dc.date | 2005-04-08 | |
| dc.date | 2008-01-05 | |
| dc.date.accessioned | 2026-07-07T08:54:55Z | |
| dc.date.available | 2026-07-07T08:54:55Z | |
| dc.description | In this paper, we aim to give a tutorial for undergraduate students studying statistical methods and/or bioinformatics. The students will learn how data visualization can help in genomic sequence analysis. Students start with a fragment of genetic text of a bacterial genome and analyze its structure. By means of principal component analysis they ``discover'' that the information in the genome is encoded by non-overlapping triplets. Next, they learn how to find gene positions. This exercise on PCA and K-Means clustering enables active study of the basic bioinformatics notions. Appendix 1 contains program listings that go along with this exercise. Appendix 2 includes 2D PCA plots of triplet usage in moving frame for a series of bacterial genomes from GC-poor to GC-rich ones. Animated 3D PCA plots are attached as separate gif files. Topology (cluster structure) and geometry (mutual positions of clusters) of these plots depends clearly on GC-content. | |
| dc.description | 18 pages, with program listings for MatLab, PCA analysis of genomes and additional animated 3D PCA plots | |
| dc.identifier | https://arxiv.org/abs/q-bio/0504013 | |
| dc.identifier | http://arxiv.org/abs/q-bio/0504013 | |
| dc.identifier | A.N. Gorban, B. Kegl, D.C. Wunsch, A. Zinovyev (eds.) Principal Manifolds for Data Visualization and Dimension Reduction, Lecture Notes in Computational Science and Engineering 58, Springer, Berlin - Heidelberg, 2008, 307-323 | |
| dc.identifier | doi:10.1007/978-3-540-73750-6_14 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/146113 | |
| dc.subject | Quantitative Methods | |
| dc.subject | Genomics | |
| dc.title | PCA and K-Means decipher genome | |
| dc.type | text |