Abstract
In this thesis, we aim to propose an interpretation of dynamic contents of video clips without any prior motion segmentation or complete motion estimation. To this end, we estimate the motion magnitudes and motion directions from the pixelwise normal flow and utilize three single Gibbs models to represent the motion distributions respectively: motion magnitude distributions along temporal domain, spatial structures of motion magnitude and spatial structures of motion direction. We measure the potential values of the three single Gibbs models by maximum likelihood criterion. In addition, in order to characterize dynamic contents in terms of the three Gibbs models, we combine the three single Gibbs models and obtain four composite Gibbs models. To demonstrate the effectiveness of the proposed models, we have applied the motion models for the application of video content classification. Experimental results show that using composite models achieves better performance than single models.