Posts

Showing posts with the label BeautifulSoup

2023-08-21: Animating Changes in Webpages: Virginia Department of Health

Image
Figure 1: Mementos from the Internet Archive's Wayback Machine showing the deletion of the entire LGBTQ youth resources page on the Virginia Department of Health website between  2023-05-31 and  2023-06-01 In July 2023, Virginia Governor Glenn Youngkin received national press for deleting webpages on the Virginia Department of Health website containing LGBTQ+ youth resources. This isn't the first time the Youngkin administration has deleted webpages: within one week of his inauguration, he deleted the Virginia Mathematics Pathways Initiative website . The Washington Post article from July 2023 also discussed additional webpages that had content removed in 2022. Journalists discovered these webpage changes through emails between state employees, but the emails didn't contain the website addresses. Without linking to mementos at a web archive, readers of the articles can't see the changes for themselves. Not linking to archived webpages in news articles is consistent w...

2018-10-19: Some tricks to parse XML files

Recently I was parsing the ACM DL metadata in XML files. I thought parsing XML is a very straightforward job provided that Python has been there for a long time with sophisticated packages such as BeautifulSoup and lxml. But I still encountered some problems and it took me quite a bit of time to figure how to handle all of them. Here, I share some tricks I learned. They are not meant to be a complete list, but the solutions are general so they can be used as the starting points to handle future XML parsing jobs. CDATA. CDATA is seen in the values of many XML fields. CDATA means Character Data. Strings inside the CDATA section are not parsed. In other words, they are kept as what they are, including marksups. One example is <script>  <![CDATA[  <message> Welcome to TutorialsPoint </message>  ]] > </script > Encoding. Encoding is a pain in text processing. The problem is that there is no way to know what the encoding the text is before ope...

2017-03-20: A survey of 5 boilerplate removal methods

Image
Fig. 1: Boilerplate removal result for  BeautifulSoup's get_text()  method for a   news website . Extracted text includes extraneous text (Junk text), HTML, Javascript, comments, and CSS text. Fig. 2: Boilerplate removal result for  NLTK's (OLD) clean_html()  method for a   news website .  Extracted text includes  e xtraneous text, but does not include Javascript, HTML, comments or CSS text. Fig. 3: Boilerplate removal result for  Justext  method for a   news website .  Extracted text includes  s maller extraneous text compared to BeautifulSoup's get_text() and NLTK's (OLD) clean_html() method, but the page title is absent. Fig. 4: Boilerplate removal result for   Python-goose  method for this   news website . No extraneous text compared to BeautifulSoup's get_text(), NLTK's (OLD) clean_html(), and Justext, but page title and first paragraph are absent. Fig. 5: Boilerplate...