Abstract
After DNA sequencing technologies were invented by Frederick Sanger from University of Cambridge in British and Walter Gilbert and Allan Maxam from Harvard University in USA in early 1970s, researchers have been able to determine DNA sequences quickly. Many huge databases including GenBank, EMBL and DDBJ have been established since 80s to store large accumulation of DNA sequences. In addition to nucleotide databases, a lot of protein databases began to be constructed such as Swiss-Prot. When Human Genome Project (HGP) was completed in 2003, biological science tends to proteomics. People started to research sequence, structure, function, and pathway of protein. Protein specific databases, analysis and predict programs are created uninterruptedly. Comparing with molecular biological databases, disease-associated databases remain fewer. Besides, most of databases are built to provide clinical information such as syndrome, diagnosis, therapeutic method and clinical care. Currently, OMIM of NCBI in USA and GeneCards of Weizmann Institute of Science in Israel can provide better disease-associated genetic information. In fact, there is not any huge database can provide useful disease-associated protein information for researchers. According to above reason, we plan to construct a new database - HDAPD that is specific to provide human disease-associated protein information. To build relationship between disease and protein is main mission of HDAPD. Next, we will collect data from several huge protein databases. Obtaining protein sequences and structures from PDB, associated pathways from KEGG and acquiring latest science literatures from PubMed. HDAPD is created to assist researchers to survey human disease- associated protein deeply and accelerate development of drugs.