Posts

Showing posts with the label GitHub

2025-02-11: Getting to the Source of the (Memento) Damage

Image
    I've previously written about the Memento Damage project, originally started by Dr. Justin Brunelle , a Web service designed to estimate the amount of damage to a web archive by assessing it's missing resources. Previously, I had been specializating some of the project  while working on the Memento Tracer project, funded by the Alfred P. Sloan Foundation , to take special considerations regarding the damage weighting for Web hosted repository pages.  I have been making further updates to the Memento Damage project over the course of this year that helps improve this analysis and damage estimation. The most prominent is the implementation of a secondary crawler component for analyzing an archived repository and its source tree. Web-hosted Git repositories are hosted on centralized Web platforms, the largest being GitHub along with other major platforms such as GitLab, Bitbucket, and Sourceforge. The source files for a Git project are hosted "behind the scen...

2019-12-21: Preserving Open Source Software with GitHub's Arctic Vault

Image
Source: Techworm GitHub is used by more than 40 million developers and currently hosts more than 100 million repositories. In early November 2019, GitHub shared plans to open the Arctic Code Vault, an effort to store and preserve open source software like Flutter and TensorFlow . With this endeavor, code for all open source projects will be stored on specialized ultra-durable 3,500-foot film with frames that include 8.8 million pixels each, designed to last 1,000 years. The data can be read by a computer or a human with a magnifying glass in case of a global power outage. "Our primary mission is to preserve open source software for future generations. We also intend the GitHub Archive Program to serve as a testament to the importance of the open source community. It’s our hope that it will, both now and in the future, further publicize the worldwide open source movement; contribute to greater adoption of open source and open data policies worldwide; and encourage long-ter...

2019-11-26: Summary of "Mentions of Security Vulnerabilities on Reddit, Twitter and GitHub"

Image
Figure 1: The Life-Cycle of a Vulnerability (Source: Horawalavithana) Cyber security attacks can be enabled by the fact that many widely-used applications share open-source libraries. As a result, a vulnerability or software weakness in one of these libraries can have far reaching impact. Once discovered, security experts may announce the vulnerability on a variety of forums, blogs, and social media sites. Cyber-adversaries  might also explore these public information channels and private discussion threads on the dark web to identify potential attack targets and ways to exploit them. In their 2019 IEEE/WIC/ACM International Conference on Web Intelligence (WI '19) paper, " Mentions of Security Vulnerabilities on Reddit, Twitter and GitHub ", Sameera Horawalavithana , Abhishek Bhattacharjee , Renhao Liu , Nazim Choudhury , Lawrence O. Hall , and Adriana Iamnitchi present a quantitative analysis of user-generated content related to security vulnerabilities on th...

2019-07-11: Raintale -- A Storytelling Tool For Web Archives

Image
My work builds upon AlNoamany's efforts to use social media storytelling to summarize web archive collections. AlNoamany employed Storify as a visualization platform. Storify is now gone . I explored alternatives to Storify in 2017 and found many of them to be insufficient for our purposes. In 2018, I developed MementoEmbed to produce surrogates for mementos and we used it in a recent research study . Surrogates summarize individual mementos. They are the building blocks of social media storytelling. Using MementoEmbed, Raintale takes surrogates to the next level, providing social media storytelling for web archives. My goal is to help web archives not only summarize their collections but promote their holdings in new ways. Raintale is the latest entry in the Dark and Stormy Archives project . Our goal is to provide research studies and tools for combining web archives and social media storytelling. Raintale provides the storytelling capability. It has been designed to vi...

2017-09-19: Carbon Dating the Web, version 4.0

Image
With this release of Carbon Date there are new features being introduced to track testing and force python standard formatting conventions. This version is dubbed Carbon Date v4.0. We've also decided to switch from MementoProxy and take advantage of the  Memgator Aggregator tool built by Sawood Alam. Of course with new APIs come new bugs that need to be addressed, such as this exception handling issue . Fortunately, the new tools being integrated into the project will allow for our team to catch and address these issues quicker than before as explained below. The previous version of this project, Carbon Date 3.0 , added Pubdate  extraction, Twitter searching, and Bing  search. We found that Bing has changed its API to only allow 30 day trials for its API with 1000 requests per month unless someone wants to pay . We also discovered a few more use cases for the Pubdate extraction by applying Pubdate to the mementos retrieved from Memgator. By default, Memgat...

2016-09-20: Carbon Dating the Web, version 3.0

Image
Due to API changes, the old carbon date tool is out of date and some modules no longer work, such as topsy . I have taken up the responsibility of maintaining and extending  the service, beginning with the following now available in Carbon Date v3.0. Carbon date 3.0 What's new New services have been added, such as bing searching , twitter searching and pubdate parsing . The new software architecture enable us to load given scripts or disable given services during runtime. The server framework has been changed from CherryPy server to  tornado server which is still a python minimalist WSGI server, with better performance. How to use the Carbon Date service Through the website , http://carbondate.cs.odu.edu : Given that carbon dating is computationally intensive, the site can only hold 50 concurrent requests, and thus the web service should be used just for small tests as a courtesy to other users. If you have the need to Carbon Date a large number of U...