pith. sign in

arxiv: 0909.2489 · v1 · pith:3JTWDIKYnew · submitted 2009-09-14 · 💻 cs.IR

PrisCrawler: A Relevance Based Crawler for Automated Data Classification from Bulletin Board

classification 💻 cs.IR
keywords bulletinboardpriscrawlersearchattachmentsrelevancesubsystemclassified
0
0 comments X
read the original abstract

Nowadays people realize that it is difficult to find information simply and quickly on the bulletin boards. In order to solve this problem, people propose the concept of bulletin board search engine. This paper describes the priscrawler system, a subsystem of the bulletin board search engine, which can automatically crawl and add the relevance to the classified attachments of the bulletin board. Priscrawler utilizes Attachrank algorithm to generate the relevance between webpages and attachments and then turns bulletin board into clear classified and associated databases, making the search for attachments greatly simplified. Moreover, it can effectively reduce the complexity of pretreatment subsystem and retrieval subsystem and improve the search precision. We provide experimental results to demonstrate the efficacy of the priscrawler.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.