Abstract
This thesis consists of 3 chapters. The first 2 chapters describe the molecular evolution of prokaryotic exceptionally large-sized genes (ELSGs) and nucleotide-sugar pyrophosphorylases/phosphorylases (NDP-sugar PPases/Pases). The third chapter describes a computer-aided drug design for searching bacterial UDP-glucose pyrophosphorylase (UGPase) inhibitors. In chapter 1, we have noted that, in contrast to the average gene size of approximately 1 kb in bacteria, 3 genes > 13 kb were present in Vibrio vulnificus. The finding prompted us to investigate the prevalence, possible function, and origin of ELSGs (>10 kb) in prokaryotes. Forty-two ELSGs (0.03%) were identified after searching more than 170,000 genes in 46 bacterial and 11 archaeal species. Homology analysis of these ELSGs indicated that, in addition to encoding non-ribosomal peptide synthetic enzymes, many ELSGs likely encode membrane-anchored proteins. Dot-matrix plot analysis of these ELSGs indicated that domain-duplication contributed significantly to size expansion. Other size expansion mechanisms were direct gene fusion, recombination of different genes, and horizontal gene transfer. In summary, ELSGs are commonly present in prokaryotes, and the evolutionary processes that have contributed to the formation of ELSGs are relatively heterogeneous. NDP-sugar PPases/Pases play a central role in providing sugar donor for the formation of glycoconjugates. Despite each of the enzymes has unique substrate specificity, they share homologous protein sequences and are therefore difficult to annotate. Chapter 2 describes our effort to better categorize these enzymes. More than 2,500 NDP-sugar PPases/Pases sequences were collected from the KEGG in this study and 86 representative sequences were selected. Phylogenetic and domain analyses revealed that members of GDP-Man PPase had the most diverse protein sequences implying that this enzyme is evolutionally closer to the common ancestor of the NDP-sugar PPases/Pases than other members of the family. Most NDP-sugar PPases/Pases were apparently derived by gene duplication although horizontal gene transfer, as in the case of eukaryotic UDP-Glc PPase, also contributed to the gene diversification. An evolutionary model for this group of enzymes was established by combining phylogenetic analysis and domain profiling. The core domains of each of the enzymes, trend of domain gain and loss, and evolutionary transition were demonstrated. These non-redundant 86 representative sequences may be used as the reference sequences for NDP-sugar PPases/Pases categorization. The low levels of sequence homology between prokaryotic and eukaryotic UGPase makes the enzyme a good target for antimicrobial drug development. Chapter 3 describes 22 UDP-Glc analogs and 97 candidate compounds, respectively selected from the ZINC and NCI compound databases using the protein docking program, LigandFit, that have potential to inhibit the bacterial UGPase activity. Five of the 97 candidate compounds were also identified by another docking program, Libdock. These 27 compounds represent potential bacterial UGPase inhibitors. The research interpreted the evolution of prokaryotic ELSGs, the categorization and evolutionary history of NDP-sugar PPases/Pases, and obtained the potential bacterial UGPases candidate inhibitors. The next step, we hope to obtain or synthesis these candidate compounds, or to screen the approved drugs in drug product database (new uses for old drugs) to perform the antimicrobial assay for the research and development of novel antimicrobial agents.