ICDAR 2021 Competition on Scientific Literature Parsing

Antonio Jimeno Yepes; Xu Zhong; Douglas Burdick

arXiv:2106.14616·cs.IR·June 29, 2021

ICDAR 2021 Competition on Scientific Literature Parsing

Antonio Jimeno Yepes, Xu Zhong, Douglas Burdick

PDF

3 Repos

TL;DR

The paper presents the results of the ICDAR 2021 Scientific Literature Parsing Competition, focusing on advancing document understanding and table recognition in scientific PDFs using datasets like PubLayNet and PubTabNet.

Contribution

It introduces a competitive benchmark for scientific literature parsing, highlighting effective methods for layout and table recognition in unstructured PDF documents.

Findings

01

High-performance object detection for layout recognition

02

Effective table component identification and post-processing

03

Impressive results enabling practical applications

Abstract

Scientific literature contain important information related to cutting-edge innovations in diverse domains. Advances in natural language processing have been driving the fast development in automated information extraction from scientific literature. However, scientific literature is often available in unstructured PDF format. While PDF is great for preserving basic visual elements, such as characters, lines, shapes, etc., on a canvas for presentation to humans, automatic processing of the PDF format by machines presents many challenges. With over 2.5 trillion PDF documents in existence, these issues are prevalent in many other important application domains as well. Our ICDAR 2021 Scientific Literature Parsing Competition (ICDAR2021-SLP) aims to drive the advances specifically in document understanding. ICDAR2021-SLP leverages the PubLayNet and PubTabNet datasets, which provide…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.