Abstract
According to the Global Health Risks Report, published by WHO, environmental issues are urged to be solved in the world. Especially, air pollution causes great damage to human health. In this work, we build an analysis framework for finding the implications between air pollution indices and cancer statistics. This framework consists of two parts for data access and data analytics, including data access flow and analytics flow, respectively. The data access flow is designed to process raw (open) data to be accessed by APIs. We map the cancer statistics to the air pollution data in the nearest monitoring stations through time and location information. The analytics flow is used to find the insights based on data exploration methods and data mining methods. The exploration methods use statistics, clustering, and series mining techniques to interpret data at hand. Then, classifiers are applied to find the relationships between air quality and cancer diseases by viewing air pollution indices and cancer statistics as features and labels, respectively. The experiments show which air pollutant has significant influence on the specific cancer. In addition, the results identified are consistent with those by traditional statistical methods. Moreover, the results achieved can also cover those by several studies. In summary, this framework is flexible and can be applied globally to other spatiotemporal data.