Logo image
Saliency-Guided Object Segmentation, Detection, and Recognition
Dissertation

Saliency-Guided Object Segmentation, Detection, and Recognition

Chang, Kai-Yueh
Doctor of Philosophy (PHD), 國立清華大學, 資訊工程學系
2011

Abstract

視覺顯著性 共同影像切割 物件偵測 物件辨識 馬可夫隨機場 圖形切割演算法 saliency co-segmentation object detection object recognition Markov random field graph cuts
The main theme of this thesis concerns the study of visual saliency and its applications to computer vision. Specifically, we strive to establish formulations that effectively utilize the saliency information to better address a number of challenging vision tasks such as object segmentation, detection and recognition. While the resulting computation models may significantly differ, they indeed are motivated by a common central idea that the saliency information is used to provide useful cues in identifying potential "regions of interest" in an image. With that, we can application-wise divide the thesis into three topics: image co-segmentation, salient object detection, and multi-class object recognition. We start with the problem of co-segmentation over multiple images, and particularly focus on two crucial issues. The first is whether a pure unsupervised algorithm can satisfactorily solve this problem. Without the user's guidance, segmenting the foregrounds implied by the common object is quite a challenging task, especially when substantial variations in the object's appearance, shape, and scale are allowed. The second issue is about the efficiency, if the technique can indeed lead to practical uses. With these in mind, we establish an MRF optimization model that has an energy function with nice properties and can be shown to effectively resolve the two aforementioned difficulties. Instead of relying on the user inputs, our approach introduces a co-saliency prior as the hint about possible foreground locations, and uses it to construct the MRF data terms. To complete the optimization framework, we design a novel global term that is more appropriate to co-segmentation and results in a submodular energy function. The proposed MRF model can thus be optimally solved by graph cuts. We demonstrate these advantages by testing our method on several benchmark datasets. Motivated by the promising results in tackling image co-segmentation, we next work on establishing a novel computational model for salient object detection. The framework explores the relatedness of objectness and saliency. It conceptually integrates the two concepts via constructing a graphical model to account for the underlying relationships, and concurrently improves their estimations by iteratively optimizing a novel energy function. Specifically, the function comprises the objectness, the saliency, and the interaction energy, respectively corresponding to explaining their individual regularities and the mutual effects. Minimizing the energy by fixing one or the other would elegantly transform the model into solving the problem of objectness or saliency estimation, while the useful information from the other concept can be utilized through the interaction term. Experimental results on two benchmark datasets demonstrate that our method can simultaneously yield a saliency map of better quality and a more meaningful objectness output for salient object detection. Finally, we show that the above framework for salient object detection can be extended to solving the problem of multi-class object recognition. In particular, we consider a practical scenario that the object colors or textures are difficult to be differentiated from the background, and the information from a depth camera is available. The knowledge about the depth can be used in learning pixel-wise classifiers for improving the quality of a bottom-up saliency map, and also in more accurately specifying the surround area of the object-level saliency estimation. To accomplish the task of object recognition, we propose a unified MRF model to simultaneously solve the segmentation and detection problems. A number of possible object segmentations (segment proposals) are linked to each superpixel in the formulation. For each superpixel, inference is carried out to derive two labels: one is the label of object class and the other is to select which proposal of object segment is most suitable if this superpixel is indeed a part of an object. Then, object recognition can be achieved by gathering information about the two types of labels from all superpixels.

Metrics

1 Record Views

Details

Logo image