Abstract
Length biased sampling has been widely used in various fields, such as epidemiology, cancer prevention trials and industrial reliability. Because of the sampling design, structure of length biased data is different to traditional survival data and the method of traditional survival analysis cannot be directly applied. Thus, we propose an estimation of semi-parametric transformation model under the right censored length biased data. The reasons to analyzed survival data based on semiparametric model is its flexibility. Besides, according to the convenience of statistical inference and reasonableness of model fitting, proportional hazards model and proportional odds models are the most commonly applied in the literature of survival analysis. However, in most situations, we need to deal with the infinite dimensional nuisance parameters. In addition, different models requires a different estimation methods, which leads inconvenience to the user on the practical. Hence, we focus on semi-parametric transformation model and propose a unified estimation procedure. Semi-parametric transformation model is a flexible model which contains proportional hazard model and proportional odd model has been widely discussed by scholars in recent years. The estimation procedure applies nonparametric maximum likelihood estimator (NPMLE) under the full likelihood in order to improve the efficiency. Besides, when the regression parameters is fixed, nuisance parameters is estimated by the self-consistency estimating equation, and we provide algorithm for implement estimation procedures. On the theoretical perspectives, we prove NPMLE provide existence, consistency property and converges to a tight Gaussian process. In simulation, we investigate the performance under different model setting, censoring rate and sample size. Compare NPMLE with methods in literature, the simulation results show that the proposed estimator provide superior efficiency. And this fact becomes more evident as sample size increasing. In the real data analysis, Alzheimer’s disease data and Channing House data are analyzed and the results of analysis has been included in the text.