Posts

Showing posts with the label digital preservation

2025-02-11: Getting to the Source of the (Memento) Damage

Image
    I've previously written about the Memento Damage project, originally started by Dr. Justin Brunelle , a Web service designed to estimate the amount of damage to a web archive by assessing it's missing resources. Previously, I had been specializating some of the project  while working on the Memento Tracer project, funded by the Alfred P. Sloan Foundation , to take special considerations regarding the damage weighting for Web hosted repository pages.  I have been making further updates to the Memento Damage project over the course of this year that helps improve this analysis and damage estimation. The most prominent is the implementation of a secondary crawler component for analyzing an archived repository and its source tree. Web-hosted Git repositories are hosted on centralized Web platforms, the largest being GitHub along with other major platforms such as GitLab, Bitbucket, and Sourceforge. The source files for a Git project are hosted "behind the scen...

2024-03-20: The Curious Case of Twitter's Server-Side UI

Image
Web archives are essential for researchers studying Twitter, offering a historical perspective on evolving conversations and trends. Archives ensure data integrity and preservation, which is crucial in a platform where content can be deleted or even, in some cases, modified. Additionally, web archives allow technical analysis of social media changes, enable comparative studies across time and platforms, and provide cultural insights, making them an essential resource for researchers in various fields. However, archiving Twitter has always been a challenge. Archiving Twitter is challenging not only due to the massive volume and dynamic nature of tweets but also because of Twitter's continually evolving user interface (UI). Studying the history of Twitter through web archives can lead to confusion, especially when encountering mementos that display different Twitter UIs from the same period.  This blog delves into the nuances of two types of Twitter UIs, client-side and server-side, ...

2023-05-25: Generative Archive Restoration

Image
  Rise of the Machines! Machine Learning just cannot seem to keep itself out of news cycles. The third version of OpenAI's generative dialogue language model, ChatGPT, had tech giants all around scrambling in quite a circus trying to push out their own versions. Google's size and stagnation in recent years had it seeing red and bringing Sergey Brin back into the fold to aid in rushing out its own chat bot, Bard . Microsoft, comparatively, has been humming along for a while now with its own research in the AI agent space but a various headlines  hint that its past and present efforts might not be paying off as well as they were hoping. You.com  is a relatively new search engine leveraging machine learning in its own chat assistant YouChat  and other services in an attempt to push the frontiers of a search engine through multi-modal search with integrated artificial intelligence enhancements. These efforts are all seeking to shake up how we seek and retrieve in...