Abstract
How to assess and quantify species compositional (dis)similarity among communities has been a central objective in ecology. Although a large number of (dis)similarity measures have been published, only the Horn index (Horn 1966), or normalized mutual information, satisfies the essential monotonicity property. Empirical estimate of the Horn index depends strongly on sample sizes and may be subject to large bias. Based on sampling data, this thesis applies Good-Turing frequency formula to derive an analytic estimator of the Horn index under any community weight proportions. Simulations results show that the proposed estimator reduces the bias associated with the empirical estimator. As the sample size increases, the proposed estimator coverages quickly to the true value. This thesis also extends the previous rarefaction and extrapolation for species richness to (dis)similarity measures which include the Sørensen index (Sørensen 1948), Horn index (Horn 1966) and Morisita index (Morisita 1959). In order to compare the (dis)similarity measures across multiple assemblages, rarefaction and extrapolation methods are proposed based on standardized sample size or sample completeness. From simulations studies, our estimated rarefaction and extrapolation curve of (dis)similarity measures can accurately quantify the species compositional (dis)similarity among assemblages up to double the sample size in each community. Biological sampling can be conducted by sampling with replacement or sampling without replacement. For the three most commonly used indices of biological diversity, including species richness, Shannon diversity, and Simpson diversity, estimators were developed for both types of sampling schemes only for species richness. In this thesis, we derive estimators of Shannon and Simpson diversities under sampling without replacement. The corresponding sample-size-based and sample-coverage-based rarefaction and extrapolation sampling curves are also developed. From simulations studies, the proposed rarefaction and extrapolation method works well up to double the sample size in each community. Real data examples are used to illustrate all proposed estimators and to demonstrate various applications.