Realistic Data Augmentation Framework for Enhancing Tabular Reasoning

Dibyakanti Kumar; Vivek Gupta; Soumya Sharma; Shuo Zhang

arXiv:2210.12795·cs.CL·October 25, 2022

Realistic Data Augmentation Framework for Enhancing Tabular Reasoning

Dibyakanti Kumar, Vivek Gupta, Soumya Sharma, Shuo Zhang

PDF

Open Access

TL;DR

This paper introduces a semi-automated data augmentation framework for tabular reasoning in natural language inference tasks, generating realistic, human-like examples to improve training, especially with limited supervision.

Contribution

The paper presents a novel semi-automated framework that creates transferable hypothesis templates and rational counterfactual tables for enhanced tabular inference data augmentation.

Findings

01

Generated human-like inference examples.

02

Improved training data quality for limited supervision.

03

Framework applicable to entity-centric tabular datasets.

Abstract

Existing approaches to constructing training data for Natural Language Inference (NLI) tasks, such as for semi-structured table reasoning, are either via crowdsourcing or fully automatic methods. However, the former is expensive and time-consuming and thus limits scale, and the latter often produces naive examples that may lack complex reasoning. This paper develops a realistic semi-automated framework for data augmentation for tabular inference. Instead of manually generating a hypothesis for each table, our methodology generates hypothesis templates transferable to similar tables. In addition, our framework entails the creation of rational counterfactual tables based on human written logical constraints and premise paraphrasing. For our case study, we use the InfoTabs, which is an entity-centric tabular inference dataset. We observed that our framework could generate human-like…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Data Quality and Management