Logo image
用類神經網路機器學習預測蛋白質之二次結構
Thesis

用類神經網路機器學習預測蛋白質之二次結構

許明智
Masters, National Tsing Hua University
1994

Abstract

類神經網路機器蛋白質資訊電腦蛋白質二次結構氨基酸錯誤後向傳遞預測準確率電腦科學 NEURAL NETWORK LEARNING ONINFORMATIONCOMPUTERProtein Secondary StructureAmino AcidNeural NetworkError Back PropagationPrediction AccuracyINFORAMTIONCOMPUTER-SCIENCE
長久以來, 在分子生物學上, 利用蛋白質的氨基酸序列來預測它的二次結構一直是一個很重要卻始終無法完全解決的問題. 雖然已經有許多預測方法被提出討論 (大致可分為四類, 即統計資訊, 樣式比對, 類神經網路,以及複合式系統), 但是這些方法的預測準確率始終停留在百分之六十到七十之間. 在本論文中, 我們對蛋白質二次結構的形成機制提出一套假設 (NDA: Neighbors Determine All 鄰居決定一切), 並且利用遞迴式類神經網路實作出一個NDA系統. 在使用實際蛋白質資料庫對這個系統加以訓練及測試後, 我們發現其實驗結果足以證實NDA和實際資料庫間有極高的相容性.也就是說, NDA可以解釋實際蛋白質資料庫中所含資料的行為. 然而因為NDA本身的限制, 這個系統無法使用於實際的預測工作中.為了利用存在此NDA系統中的資訊, 我們另外提出了兩種方法: 單邊NDA和NECN. 單邊NDA簡化了NDA的嚴格限制, 因此可以於特定的實際預測工作中使用. 而NECN則是結合了NDA網路以及傳統網路架構. 先從NDA網路中抽取出新的氨基酸代碼, 再將這代碼應用於傳統的網路架構上, 形成一個兩階段式的架構. 我們使用實際蛋白質資料庫以及人造NDA資料庫 (根據NDA假設造出一個人造資料庫) 來訓練及測試NECN後, 得到的結果證實了NECN的確優於傳統使用的網路架構.The prediction of the secondary structure of a protein from itsamino acids sequence has long been a not-yet-completely solvedproblem despite its importance in molecular biology. Althoughvarious prediction methods have been proposed, the predictionaccuracy still hovered around 60-70%. In this thesis weproposed a hypothesis on the formation mechanism of secondarystructures called NDA ( Neighbors Determine All ). A NDAimplementation using recurrent neural network is tested on adatabase consists of proteins with known secondary structuressequence. The result confirms that NDA hypothesis is highlycompatible to the database. However, the NDA network cannot beapplied on practical prediction tasks due to certain reasons.To use the information contained in NDA network we proposed twoapproaches: the Single-Side NDA and NECN. The Single-Side NDArelieves the strict restrictions of NDA and providespossibility of practical application. The NECN ( NDA ExtractedCode Neural network ) combines a NDA neural network and aneural network which is traditionally used in this problem toform a two phase secondary structure prediction system. NECN istested with the protein database and another databaseconstructed using NDA hypothesis. Both results show that NECNperforms better than conventional neural network methods usedin this problem. Although the performance gain is not large,the NECN can still be improved in many aspect. In addition, theNDA hypothesis and its neural network implementation are alsouseful tools in this task.

Metrics

1 Record Views

Details

Logo image