Logo image
Clustering for Web information hierarchy mining
Conference paper

Clustering for Web information hierarchy mining

Hung-Yu Kao, Ming-Syan Chen and Jan-Ming Ho
Proceedings - IEEE/WIC International Conference on Web Intelligence, WI 2003, pp.698-701
2003

Abstract

Assembly Clustering algorithms Data mining Electronic mail HTML Information science Joining processes Merging Scalability Web pages Artificial Intelligence Information Systems Computer Networks and Communications Human-Computer Interaction Information Systems and Management
Benefiting from the growth of techniques of dynamic page generation, the amount and the complexity of Web pages increase explosively. The structures of Web pages which are dynamically generated by the same templates are thus similar to one another and are usually assembled by a set of fundamental information clusters These neighboring information clusters usually represent the similar semantics and form a larger cluster with the more generalized information. The hierarchical structure generated by information clusters in a bottom-up manner is called the information hierarchy of a page. We study the problem of mining the information hierarchies of pages in Web sites to recognize the information distribution of pages within the multilevel, multigranularity configurations. Explicitly, we propose an information clustering system that applies a top-down information centroid searching algorithm and a multigranularity centroid converging process on the document object model (DOM) trees of pages to build the information hierarchies of pages. Experiments on several real news Web sites show the high precision and recall rates of the proposed method on determining information clusters of pages and also validate its practical applicability to real Web sites.

Metrics

1 Record Views

Details

Logo image