Scinovex
article Open AccessTop 10% cited

Identification and application of the concepts important for accurate and reliable protein secondary structure prediction

Protein Science · 1996 · Vol. 5(11) · pp. 2298–2310
Ross D. KingMichael J.E. Sternberg

Abstract

A protein secondary structure prediction method from multiply aligned homologous sequences is presented with an overall per residue three-state accuracy of 70.1%. There are two aims: to obtain high accuracy by identification of a set of concepts important for prediction followed by use of linear statistics; and to provide insight into the folding process. The important concepts in secondary structure prediction are identified as: residue conformational propensities, sequence edge effects, moments of hydrophobicity, position of insertions and deletions in aligned homologous sequence, moments of conservation, auto-correlation, residue ratios, secondary structure feedback effects, and filtering. Explicit use of edge effects, moments of conservation, and auto-correlation are new to this paper. The relative importance of the concepts used in prediction was analyzed by stepwise addition of information and examination of weights in the discrimination function. The simple and explicit structure of the prediction allows the method to be reimplemented easily. The accuracy of a prediction is predictable a priori. This permits evaluation of the utility of the prediction: 10% of the chains predicted were identified correctly as having a mean accuracy of > 80%. Existing high-accuracy prediction methods are "black-box" predictors based on complex nonlinear statistics (e.g., neural networks in PHD: Rost & Sander, 1993a). For medium- to short-length chains (> or = 90 residues and < 170 residues), the prediction method is significantly more accurate (P < 0.01) than the PHD algorithm (probably the most commonly used algorithm). In combination with the PHD, an algorithm is formed that is significantly more accurate than either method, with an estimated overall three-state accuracy of 72.4%, the highest accuracy reported for any prediction method.

Protein Structure and DynamicsEnzyme Structure and FunctionMachine Learning in BioinformaticsProtein secondary structureAlgorithmComputer scienceA priori and a posterioriSequence (biology)Protein structure predictionNonlinear systemMathematicsProtein structureChemistry

MeSH terms

Amino Acid SequenceCalorimetry, Differential ScanningModels, ChemicalMolecular Sequence DataProtein Structure, Secondary
Citations
459
FWCI
8.10
field-weighted impact
References
59
Percentile
98%
vs. same field & year
Citations per year
Cited by
Application of multiple sequence alignment profiles to improve protein secondary structure prediction
Proteins Structure Function and Bioinformatics · 2000 · 789 citations
Evaluation and improvement of multiple sequence methods for protein secondary structure prediction
Proteins Structure Function and Bioinformatics · 1999 · 690 citations
NPS@: Network Protein Sequence Analysis
Trends in Biochemical Sciences · 2000 · 1,744 citations
References
Prediction of Protein Secondary Structure at Better than 70% Accuracy
Journal of Molecular Biology · 1993 · 2,942 citations
Prediction of protein conformation
Biochemistry · 1974 · 3,338 citations
Database of homology‐derived protein structures and the structural meaning of sequence alignment
Proteins Structure Function and Bioinformatics · 1991 · 1,658 citations
Classification and regression trees
European Journal of Operational Research · 1985 · 10,158 citations
Classification and Regression Trees.
Journal of the American Statistical Association · 1986 · 21,013 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.