Abstract
In large scale assessment programs, the method of booklet design is commonly adopted for sampling a large amount of test items and describing latent traits of student participants. The missingness derived from booklet design not only reduce the number of samples responding to each tested item, but also cause vacancy in participants’ responses, which in turn harms the result of DIF assessment. Recently, an increasing attention has been drawn to a non-parametric DIF assessment method called the Mantel-Haenszel test (MH) due to its simplicity. Although past studies have found the matching strategy was crucial to the effectiveness of MH-DIF assessment in tests adopting booklet design, researches comparing the effectiveness of various matching strategies developed recently within a more authentic but complex booklet design context areis relatively rare. Furthermore, the DIF assessments in previous studies are often conducted based on a limited, or even without, scale purification procedure, and the differences of mean difficulties between booklets are too small to influence the result of DIF assessment. In this research, three research questions were raised. First, can we amplify the differences in MH-DIF results using equated pooled booklet matching strategy and the other matching strategies by increasing the difference of mean difficulty between booklets? Second, if the matching variable is iteratively purified according to presumed results of DIF assessment instead of true DIF items, will this purification procedure affect the performance of DIF assessment among various matching strategies? Third, what is the difference in the results of DIF assessments between various matching strategies in both main and sub-dimension of PISA 2012 booklet design, respectively? In this researchstudy, the sample size, impact, percentage of DIF items, and the range of mean item difficulty between booklets are manipulated. The booklet design adopted in this study research followed authentic PISA 2012, and we the matching variable based on presumed DIF assessment is iteratively purified matching variable based on presumed DIF assessment. Type I error rate and power of DIF assessment using three matching strategies, namely block level, percent pooled booklet, and equated pooled booklet, are recorded. The findings indicated that the power rates in sub-dimension were higher than that those in main dimension among three matching strategies. Controlling for dimension and sample size, the power rates of block level matching strategies became lower when there were more DIF items in the test. The power rates of percent pooled booklet were affected by the three way interaction of impact, range of mean item difficulty between booklets, and percentage of DIF items. Comparing to the previous two strategies, equated pooled booklet strategy yielded the most ideal Type I error rate and the highest power rate in all scenarios. Furthermore, a real data example derived from PISA 2012 math test of Taiwan was analyzed for gender DIF using the equated pooled booklet strategy. Approximately 30% of the items were deemed to be DIF. According to this research, equated pooled booklet strategy with iterative purification procedure is strongly recommended in DIF assessment, especially when there are huge differences of mean difficulty between booklets, or when a lot of DIF items are expected in tests.