Posts

Showing posts with the label Table Recognition

2023-01-10: A Summary of "Multi-Type-TD-TSR -- Extracting Tables from Document Images using a Multi-stage Pipeline for Table Detection and Table Structure Recognition: from OCR to Structured Table Representations"

Image
                                              Figure 1: Detecting tables and extracting table cell structures in document image (Fig 2 in  Fischer et al. ) In the past decades, several works have been published on detecting and extracting tables both from in-text-tables and  tables appearing in born-digital or scanned PDF documents  [ Pyreddi et al. ]. Early work focus on using heuristics such as character alignment in table images to extract tables [ Pyreddi et al. ]. Recent works involve detecting the corners of the table cells and inferring their connectivity [ Seo et al. ] in document images (see Figure 2).                                                           Figure 2: Detecting table cell corners (Fig 2 in...

2022-12-29: A Summary of "CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documents"

Image
                                                                                Figure 1: Traditional  and Deep Learning Approaches to Table Recognition ( Hashmi et al. ) Table recognition refers to the process of using optical character recognition (OCR) and machine learning (ML) models to identify the rows, columns, and individual text cells in tables in digital documents either born-digital or scan PDFs. The task of table recognition has been under investigation for more than two decades  for automatically extracting textual information from a variety of tables  [ Kieninger et al. , Wei et al. ].  Automatic table recognition can be very challenging due to tables having different structures, data types, and misaligned data e...