Deutsch
 
Hilfe Datenschutzhinweis Impressum
  DetailsucheBrowse

Datensatz

DATENSATZ AKTIONENEXPORT

Freigegeben

Hochschulschrift

Automatic Generation of Thematically Focused Information Portals from Web Data

MPG-Autoren
/persons/resource/persons45500

Sizov,  Sergej
Databases and Information Systems, MPI for Informatics, Max Planck Society;
International Max Planck Research School, MPI for Informatics, Max Planck Society;

Volltexte (beschränkter Zugriff)
Für Ihren IP-Bereich sind aktuell keine Volltexte freigegeben.
Volltexte (frei zugänglich)
Es sind keine frei zugänglichen Volltexte in PuRe verfügbar
Ergänzendes Material (frei zugänglich)
Es sind keine frei zugänglichen Ergänzenden Materialien verfügbar
Zitation

Sizov, S. (2005). Automatic Generation of Thematically Focused Information Portals from Web Data. PhD Thesis, Universität des Saarlandes, Saarbrücken. doi:10.22028/D291-23767.


Zitierlink: https://hdl.handle.net/11858/00-001M-0000-000F-24F5-4
Zusammenfassung
Finding the desired information on the Web is often a hard and time-consuming task. This thesis presents the
methodology of automatic generation of thematically focused portals from Web data. The key component of the proposed
Web retrieval framework is the thematically focused Web crawler that is interested only in a specific, typically small, set of topics. The focused crawler uses classification methods for filtering of fetched documents and identifying most likely relevant Web sources for further downloads.

We show that the human efforts for preparation of the focused crawl can be minimized by automatic extending of the training dataset using additional training samples
coined archetypes. This thesis introduces the combining of classification results and link-based authority ranking methods for selecting archetypes, combined with periodical re-training of the classifier. We also explain the architecture of the focused Web retrieval framework and discuss results of comprehensive use-case studies
and evaluations with a prototype system BINGO!.

Furthermore, the thesis addresses aspects of crawl postprocessing, such as refinements of the topic structure and restrictive document filtering. We introduce postprocessing methods and meta methods that are applied in an restrictive manner, i.e. by leaving out some uncertain documents rather than assigning them to inappropriate topics or clusters with low confidence. We also introduce the methodology of collaborative crawl postprocessing for
multiple cooperating users in a distributed environment, such as a peer-to-peer overlay network.

An important aspect of the thematically focused Web portal is the ranking of search results. This thesis addresses the aspect of search personalization by aggregating explicit or implicit feedback from multiple users and capturing topic-specific search patterns by profiles. Furthermore, we consider advanced link-based authority ranking algorithms that exploit the crawl-specific information, such as classification confidence grades for particular documents.
This goal is achieved by weighting of edges in the link graph of the crawl and by adding virtual links between highly relevant documents of the topic.

The results of our systematic evaluation on multiple reference collections and real Web data show the viability of the proposed methodology.