Posts

Showing posts with the label LANL

2022-12-21: PDFServer - Our Summer Internship at LANL

Image
This summer (summer of 2022), we (Himarsha Jayanetti, Gavindya Jayawardena, and Yasith Jayawardana) got the opportunity to intern at the Los Alamos National Laboratory (LANL) under the Graduate Research Assistant Program . Himarsha and Yasith worked as a part of the Library's Research and Prototyping Team  (Proto Team) under the supervision of Drs. Martin Klein and Shawn M. Jones . Gavindya was supervised by Brian Cain on the Institutional Scientific Content (ISC) Team. Each of us was responsible for one of the three main components that make up the "PDFServer". What is the PDFServer? PDFServer is a prototype to support LANL's Review and Approval System for Scientific and Technical Information (RASSTI) , which is the starting point for collecting LANL’s documented scientific output and it provides workflows for authors and reviewers. The overall architecture of the PDFServer is displayed in Figure 1. We selected Laravel 9.11 , which is a PHP web application framewo...

2022-02-23: One in Five arXiv Articles Reference GitHub

Image
Starting in Fall 2021, I've had the opportunity to work on the  CoSAI Project  under the guidance of  Dr. Martin Klein ,  Dr. Michael Nelson , and  Dr. Michele Weigle . The CoSAI Project is working to preserve web-based scholarship including source code. The goal of the project is to make the archival process more accessible to institutions by creating a curation workflow to facilitate the process. As part of the project, we wanted to find a set of code repository URIs that were referenced in scholarly publications. To do this, we decided to extract URIs from PDFs in the  arXiv  corpus which now includes more than 2 million papers . We focused on a corpus of 1.56 million PDFs from April 2007 to November 2021. During an internship at LANL in Summer 2021 , Yasith Jayawardana created code that Robustifies URIs found in PDFs. Part of the code extracts URIs found in PDFs using the PyPDFium2 and PyPDF2 to extract annotated URIs and URIs in the text, respec...