Posts

Showing posts with the label unarchived web

2018-07-18: Why We Need Private Web Archives: Almost Two-Thirds of Web Traffic IS NOT Publicly Archivable

Image
Google.com mementos from May 8th 1999 on the Internet Archive In terms of the ability to be archived in public web archives, web pages fall into one of two categories: publicly archivable, or not publicly archivable. 1. Publicly Archivable Web Pages: These pages are archivable by public archives. The pages can be accessed without login/authentication. In other words, these pages do not reside behind a paywall. Grant Atkins examined paywalls in the Internet Archive for news sites and found that web pages behind paywalls may actually be redirecting to a login page at crawl time. A good example of a publicly archivable page is   Dr. Steven Zeil's page  since no authentication is required to view the page. Furthermore, it does not use client-side scripts (i.e., Ajax) to load additional content, so what you see in the web browser and what you can replay from public web archives are exactly the same. Screen shot from  Dr. Steven Zeil's page  capture...

2016-10-03: Summary of “Finding Pages on the Unarchived Web"

Image
by: Hugo C. Huurdeman , Anat Ben-David , Jaap Kamps , Thaer Samar , and Arjen P. de Vries Proceedings of the 14th ACM/IEEE-CS Joint Conference on Digital Libraries 2014 In this paper , the authors detailed their approach to recover the unarchived Web based on links and anchors of crawled pages. The data used was from the Dutch 2012 Web archive at the National Library of the Netherlands (KB) , totaling about 38 million webpages. The collection was selected by the library based on categories related to Dutch history, social and cultural heritage. Each website is categorized using UNESCO code . The authors try to address three research questions: Can we recover a significant fraction of unarchived pages?, How rich are the representations for the unarchived pages?, and Are these representations rich enough to characterize the content? The link extraction used Hadoop MapReduce and Apache Pig to process all archived webpages and used JSoup to extract links from their content. ...