Posts

Showing posts with the label browser based preservation

2026-02-13: Paper Summary: "High Fidelity Web Archiving of News Sites and New Media with Browsertrix"

Image
Figure 1: The Browsertrix Tool Suite For the Saving Ads project and the Game Walkthroughs and Web Archiving project , we have used Browsertrix Crawler , ArchiveWeb.page , and ReplayWeb.page , which are Webrecorder tools that were discussed in Walsh et al.’s paper “ High Fidelity Web Archiving of News Sites and New Media with Browsertrix .” In this paper, Walsh et al. describe tools that are integrated with Browsertrix and the features that differentiated their tools from other web archive crawlers and replay systems. Browsertrix is a free and open-source web archiving platform that can be run locally , self-hosted , or used through Webrecorder's hosted service . Browsertrix uses Browsertrix Crawler for archiving web pages, ArchiveWeb.page to patch archived web pages, and ReplayWeb.page to replay archived web pages (Figure 1). Walsh, Tessa, Henry Wilkinson, and Ilya Kreymer. “High Fidelity Web Archiving of News Sites and New Media with Browsertrix.” in Proceedings of the 2024 Int...

2024-01-19: Paper Summary of "Overcoming Barriers to Information Exchange on the Web" by Ayush Goel

Image
Fig. 1: Figure 7 from " Sprinter: Speeding Up High-Fidelity Crawling of the Modern Web " Introduction      Dr. Ayush Goel is a recent Ph.D graduate from the University of Michigan’s Systems Lab , studying under Dr. Harsha V. Madhyastha and Dr. Ravi Netravali , now working at Hewlett-Packard Labs as a Research Scientist. While the UMich Systems Lab does not necessarily work on Web archiving, much of their work has overlap with our own here at the Web Science and Digital Libraries group. I first became aware of Dr. Goel’s work through his previous work and publication, “ Jawa: Web Archival in the Era of JavaScript ”, which was mentioned by former WS-DL alumnus Emily Escamilla in her 2022 IIPC Web Archiving Conference trip report , and again in her 2023 IIPC WAC trip report . The Jawa project is an impressive accomplishment on its own but it is only a piece of Dr. Goel’s contributed body of work. With impacts on my own research work and long-lasting impacts for the We...

2018-12-03: Acidic Regression of WebSatchel

Image
Mat Kelly reviews WebSatchel, a browser based personal preservation tool.                                                                                                                                                               ...

2018-04-30: A High Fidelity MS Thesis, To Relive The Web: A Framework For The Transformation And Archival Replay Of Web Pages

Image
It is hard to believe that the time has come for me to write a wrap up blog about the adventure that was my Masters Degree and the thesis that got me to this point. If you follow this blog with any regularity you may remember two posts, written by myself, that were the genesis of my thesis topic: 2017-01-20: CNN.com has been unarchivable since November 1st, 2016 2017-03-09: A State Of Replay or Location, Location, Location Bonus points if you can guess the general topic of the thesis from the titles of those two blog posts. However, it is ok if you can not as I will give an oh so brief TL;DR;. The replay problems with cnn.com were, sadly, your typical here today gone tomorrow replay issues involving this little thing, that I have come to , known as JavaScript. What we also found out, when replaying mementos of cnn.com from the major web archi...