Logo image
網頁篩選的索引結構
Thesis

網頁篩選的索引結構

陳珍珮
Masters, National Tsing Hua University
1997

Abstract

索引網址資訊篩選 WWWindexURLInformation FilteringWWW
Web 風潮帶領了網頁的快速成長,使得如何有效率地搜尋與取用 Web 上的資訊成為一項重要議題。使用者必須利用 web 上資訊檢索的服務,像是搜尋引 擎(如 HotBot, Lycos)和 web 目錄(如 Yahoo),來從廣大的 web 世界中找到資訊。但目前的普遍作法,即以關鍵字為主的搜尋,很明顯地並不足以描述使用者需求,因為這樣的查詢往往帶來仍然相當大量而顯得意義模糊的網頁作為比對結果。為了能要更豐富查詢的意涵,並且因此縮減搜尋空間,應該利用除了關鍵字之外的網頁相關資訊。URL(網址)就是一項很好的候選者,因 為每個網頁都會有(refer to)一個網址,並且就算使用者並不真正清楚 "到哪裡去找什麼" (where to findwhat),仍能夠很輕易地給定一個網址查詢。舉例來說,想要查詢代理人"Agent" 相關資訊,可以下列的網址查詢將查詢 空間限制在美國各大學的網址,或甚至指定是其中的資訊系所:"cs.*.edu"。利用網址查詢,使用者可以將他們的搜尋限制在一群可信賴的,或者有趣的來源內。除了單次查詢,我們預期很快地 Web 搜尋服務將會利用資訊篩選的技術將新登記的網頁通告給使用者,而非只是在查詢時將他們排序列在數百萬的其他網頁間。資訊篩選已經廣泛地被使用在電子郵件和新聞論壇的服務上。藉著指定描述檔,使用者會定期地收到和他在描述檔中指定的關鍵字相關的新的網頁的消息。另一方面,藉著這項服務,新的網頁能夠較容易被瀏覽,而非被數以萬計的相似網頁所掩埋。根據上述的觀察,我們提供了數種建立描述檔索引的方法,以及對網址和大量描述檔做比對的演算法。我們也提供了模擬的實驗結果,來比較他們在各種不同設定下的表現。The rapid growth of web pages makes searching and retrievinginformation on web efficiently and effectively a criticalproblem. While common keyword-based search engines bringvague and numerous results, we introduce the concept toquery against. URLs which would bound the search results into amuch smaller fraction. As conventional informationfiltering services, in which users submit profiles such thatthey will be automatically informed of new addtions that maybe of intrest, we expect URL profiles will soon be applied onsimilar applications.In this thesis we propose several index structures for indexingprofiles and algorithms that efficiently match URLs against alarge number of profiles. We also present simulation resultsto compare their performance under different scenarios.

Metrics

1 Record Views

Details

Logo image