Abstract
More recently, technological advances have made itpossible to process large amounts of image data; the mainviewpoint of these developed techniques is generally based on the engineering viewpoint. That is, the scheme of these techniques is focused on the quality of results and the development of algorithms. Additionally, one important viewpoint called human vision mechanism is also usually to be used in the recent researches; hence, human vision mechanism is significant information. In order to further realize their potential in the application of image processing, a system for the viewpoint of human-vision base needs to be developed.Based on human-vision base properties, the top-downprocess and bottom-up process are usually adopted techniques andapplied to diverse researches. In our proposed system, we utilize these properties to process from low-level scene analysis for local image properties to high-level scene analysis for semantic description.For bottom-up process, a image segmentation algorithm isproposed based on Self-Organizing Map (SOM) methodology whichtakes into account the color similarity and spatial relationships of objects within an image. Based on the features of color similarity, an image is first segmented into coarse cluster regions, named planes, using SOM_1 algorithm with a labeling process. The final segmented regions are treated by computing the spatial distance between any two planes and using SOM_2 algorithm with a labeling process. Moreover, the selection of parameters, named the number of iterations and output nodes for SOM algorithm is also discussed in this approach. The segmented objects, which are similar to human perceived, are represented for the proposed approach. Experiments show that this approach is reliable and feasible. It can provide the primary information to further investigate the image content.For top-down process, we have also investigated the imagecontent descriptions which adopt the image features and spatialinformation. A forward-recall image processing system has beenproposed which contains the forward process with a semanticdescription for each segmented object and the recall process with redrawing each of segmented objects based on the semanticdescription and spatial location. The forward process involvingthe bottom-up and top-down processes is represented the semanticinterpretation obtained using the features of color, texture,spatial relationships that represent to the indexes, and thenusing these indexes to construct the linguistic inference ruledatabase based on human experiences and knowledge. Thecorresponding features for each object in an image can beobtained. Furthermore, each region can be represented itscorresponding interpretation by operating the inference ruledecision in linguistic data base. Each region, in general, can be almost exacted to interpret its semantic meaning description in our experiments.Through our researches mentioned above, a human-visionbase image understanding system containing image segmentation,linguistic meaning interpretation and recall is proposed.