Developing optimal nonlinear scoring function for protein design
| dc.creator | Hu, Changyu | |
| dc.creator | Li, Xiang | |
| dc.creator | Liang, Jie | |
| dc.date | 2004-07-29 | |
| dc.date.accessioned | 2026-07-07T05:58:34Z | |
| dc.date.available | 2026-07-07T05:58:34Z | |
| dc.description | Motivation. Protein design aims to identify sequences compatible with a given protein fold but incompatible to any alternative folds. To select the correct sequences and to guide the search process, a design scoring function is critically important. Such a scoring function should be able to characterize the global fitness landscape of many proteins simultaneously. Results. To find optimal design scoring functions, we introduce two geometric views and propose a formulation using mixture of nonlinear Gaussian kernel functions. We aim to solve a simplified protein sequence design problem. Our goal is to distinguish each native sequence for a major portion of representative protein structures from a large number of alternative decoy sequences, each a fragment from proteins of different fold. Our scoring function discriminate perfectly a set of 440 native proteins from 14 million sequence decoys. We show that no linear scoring function can succeed in this task. In a blind test of unrelated proteins, our scoring function misclassfies only 13 native proteins out of 194. This compares favorably with about 3-4 times more misclassifications when optimal linear functions reported in literature are used. We also discuss how to develop protein folding scoring function. | |
| dc.description | 25 pages, 6 figures, 7 tables. Accepted by Bioinformatics | |
| dc.identifier | https://arxiv.org/abs/q-bio/0407040 | |
| dc.identifier | http://arxiv.org/abs/q-bio/0407040 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/88406 | |
| dc.subject | Biomolecules | |
| dc.title | Developing optimal nonlinear scoring function for protein design | |
| dc.type | text |