pith. sign in

arxiv: 1612.05413 · v1 · pith:3TW36FTSnew · submitted 2016-12-16 · 💻 cs.DL · cs.IR

Analyzing Web Archives Through Topic and Event Focused Sub-collections

classification 💻 cs.DL cs.IR
keywords archiveschallengessub-collectionsfocusedworkinganalyzingapproacharchive
0
0 comments X
read the original abstract

Web archives capture the history of the Web and are therefore an important source to study how societal developments have been reflected on the Web. However, the large size of Web archives and their temporal nature pose many challenges to researchers interested in working with these collections. In this work, we describe the challenges of working with Web archives and propose the research methodology of extracting and studying sub-collections of the archive focused on specific topics and events. We discuss the opportunities and challenges of this approach and suggest a framework for creating sub-collections.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.