Posts

2024-10-07: Leveraging LLMs for Transcript Generator - Summer Internship Experience at Amazon Inc.

Image
  This summer, I had the privilege of interning at Amazon with the Alexa Certification Technology (Alexa-Cert-Tech) team located in Sunnyvale, California, USA. My internship was a 13-week program which started on May 28th, 2024. During this internship, I worked mostly as a data scientist intern under the supervision of Praveen Chinnusamy , Bheema Rajulu , and Jingya Li . Throughout this program, I attended weekly meetings with the entire Alexa-Cert-Tech team. The weekly meetings were to update my progress, obtain feedback, resolve issues, or improve the solution. I had one-on-one meetings with my mentors Bheema Rajulu and Jingya Li twice weekly to discuss my progress and any issues I faced. I also met with my manager Praveen Chinnusamy at least once a week. My project was focused on automating Behavior-Driven Development (BDD) test statement generation and improving transcript generation using Large Language Models (LLMs) . This blog post will dive deep into the challenges, s...

2024-10-03: The End of a Long Journey: Reflecting on My PhD Experience

Image
In the Spring of 2019, I embarked on a six-year journey toward earning my PhD in Computer Science. After completing my undergraduate degree in Computer Science and Engineering at the University of Moratuwa , Sri Lanka, in 2018, I applied to the Computer Science PhD program at Old Dominion University (ODU). I was fortunate to be accepted to the Neuro Information Retrieval and Data Science (NIRDS) Lab , which is part of ODU’s Web Science and Digital Libraries (WSDL) Research Group . Led by Dr. Sampath Jayarathna , the NIRDS lab focuses on applied research involving human subjects and multimodal biosignals. Early in my PhD, before defining my research focus, I explored how to operate our lab's various biosignal acquisition devices and software (Eye Trackers, EEG Devices, and Wearable Health Sensors). This exploration allowed me to gain a deep understanding of the tools and technologies around human-subjects research, as well as their capabilities and limitations. A recurring issue th...

2024-10-02: MS Thesis: Surfacing Text Changes in Archived Webpages

Image
Thesis defense, July 29, 2024. Picture courtesy of Dr. Michele Weigle. My master’s thesis, “Surfacing Text Changes in Archived Webpages” explores how users can better find and view changes on webpages in web archives. The thesis contributes to the area of information seeking behavior in web archives, and addressed three research questions.  1. How can we make changes in webpages discoverable and understandable? We presented a change text search interface for web archives that allows users to find changes in webpages. This interface also includes an animated deletion tool and a sliding difference tool, which help users view the changes in context. This part of the thesis was informed by our formative investigation “ User Tasks of Journalists .” We presented this work in our paper, “ Making Changes in Webpages Discoverable: A Change-Text Search Interface for Web Archives ” at JCDL 2023 , and the paper earned the best student paper award. 2. How can we increase efficiency in web archi...

2024-09-22: Looking Back at My PhD Journey

Image
I ( Gavindya Jayawardena ) began my PhD journey in Spring 2019 under the guidance of Dr. Sampath Jayarathna at Old Dominion University (ODU) , immediately after completing my Bachelor’s degree. I joined the Neuro-Information Retrieval and Data Science (NIRDS) lab , a subgroup within the Web Science and Digital Libraries (WSDL) research lab in the Computer Science Department . Dr. Sampath Jayarathna provided me with the opportunity to engage in eye-tracking research, which rapidly became the central focus of my studies. In collaboration with Dr. Anne Perrotti from ODU, I explored how eye-tracking measurements could be utilized to predict Attention-Deficit/Hyperactivity Disorder (ADHD) using machine learning algorithms. By extracting raw eye-tracking data, I created a feature set that enabled multiple models to achieve high prediction accuracy with tree-based classifiers. Additionally, I studied the performance of adolescents with ADHD during an audiovisual Speech-In-Noise t...

2024-09-20: Some URLs Are Immortal, Most Are Ephemeral

Image
This post reports preliminary results from the "Not Your Parents' Web" project, a collaboration between Old Dominion University's Web Science and Digital Libraries group , the Internet Archive (IA), and the Filecoin Foundation , with funding provided by the Filecoin Foundation. This work was performed by Kritika Garg (ODU PhD student), Sawood Alam (Internet Archive), Michele Weigle (ODU faculty), Michael Nelson (ODU faculty), and Dietrich Ayala (Filecoin Foundation). Our goal is to revisit the question, "How long does a webpage last?". The canonical response has been 44 , 75 , or 100 days, but that was based on research done in the early days of the Web (1996–2003). A study published in May 2024 from the Pew Research Center, "When Online Content Disappears" , found that 38% of webpages that existed in 2013 are no longer available on the live Web. A February 2024 study from the SEO company Ahrefs also considered the current status of links from 2...

2024-09-13: Paper Summary: Uncertainty Quantification in Table Structure Recognition

Image
                   Figure 1. An illustration of the differences between aleatoric and epistemic uncertainties (Yang et al., 2023). Introduction Table Structure Recognition (TSR) is a task of document analysis that focuses on identifying rows and columns in digital table images [4]. While current TSR methods can identify cell locations, they lack the ability to predict uncertainties in their results [1]. This limitation has hindered the real-world application of TSR, such as automatically extracting data from table images in physical sciences. In this blog post, we summarize our paper titled " Uncertainty Quantification (UQ) for Table Structure Recognition ", presented at the 2024 IEEE International Conference on Information Reuse and Integration for Data Science . In this paper, we proposed a method called TTA-m (Test-Time Augmentation with multiple models) that aims to quantify uncertainties in TSR predictions, potentially enhancing ho...