Abstract
MicroRNAs are a class of small non-coding RNAs which play important regulatory roles in animals and plants. They cause transcriptional cleavage or translational repression through binding their target mRNAs. MicroRNAs affect a variety of cellular processes such as development, cell proliferation, apoptosis, and stress response. Thus identification of mRNA targets is an essential step to understand microRNA functions. Currently several microRNA target prediction tools have been developed. The majority of these algorithms are based on the sequence alignment or the minimum free energy of the hybridization. However, due to the omission of gene expression information in the screening process, a number of candidate targets could be false positives which are too large to validate. In this work, a filtering strategy which was implemented based on SVM machine learning has been built in order to reduce the false positives. Information of sequence alignment retrieved from existing database and microarray expression data were both used to classify mRNA candidates into non-target or target group by SVM, trying to separate microRNA target genes from non-target genes. Besides, the concept of conservation between species has been included to mitigate the problem of noisy data then decrease false positive predictions.