Abstract
The researches of active site predictions that are based on the analysis of sequence and structure are increasingly developing for the past few years. The scope of enzyme categories in prediction is no longer limited to some specific enzyme families. In this thesis, we provide a partial least squares regression model trained over 225 nonhomologous enzymes to predict the location of actives sites. We use conservation scores, catalytic propensity, intrinsic dynamics of enzymes, relative solvent accessibility(RSA), pKa changes, the average RSA deviation in sequential residues, distances between residues and domain center and the spatial clustering scores as prediction model inputs. The performance of our predictions is interpreted by sensitivity=0.35, specificity=0.54 and Matthews correlation coefficient(MCC)=0.38 when we select the top 2 candidates. Sensitivity=0.66, specificity=0.29 and MCC=0.38 when we select the top 7 candidates. The dominant features of residues of enzymes in PLS model are conservation, spatial clustering scores of prediction candidates, distance between residues and domain center and the average RSA deviation in sequential residues.