Abstract
Abstract With rapidly generated whole genome sequence data especially those microbial organisms, it can be used to explore the diversity of ancient life. More and more comparative genomic methods have been used to investigate the similarities or dissimilarities between organisms. Phylogenetic tree based on 16S rRNA indicates the prokaryotic evolutionary relationship unrevealed from the morphological characteristics. Other features like GC content and amino acid composition are widely used to account for extreme environmental organisms such as thermophiles, acidophiles, halophiles, etc. Quarrying the whole genome wide information may suggest why microbe diverse. Here we constructed a database GPDB (Genome Profile DataBase) with 145 microbial genomes including bacteria and archaea. The original sequence data and annotations are based on NCBI GeneBank and RefSeq databases. The uniform nomenclature and classification were used according to the taxonomy database at NCBI. In order to automatically process so many features, the program called "Genome Profile Pipeline" has been developed in perl language. Here we present lots of various "Genome Profile", such as basic information (taxonomy, genome size, orf number…), nucleotide composition (GC & AT content, total GC & AT skew, N-nucleotide frequency, codon usage…), and amino acid composition (N-peptide frequency distribution, proteome distribution like length & Mw & pI & transmembrane helix protein & fold…) in graphic ways. In order to estimate different combination interactively, an on-line graphic browsing interface which use Euclidean distance for hierarchical clustering method was built to compare and view the difference between these organisms. Further more, the website is modulated for more Genome Profile to be included and compared in the future.