Denoising Pre-Training and Data Augmentation Strategies for Enhanced RDF   Verbalization with Transformers

Sebastien Montella; Betty Fabre; Tanguy Urvoy; Johannes Heinecke; Lina; Rojas-Barahona

arXiv:2012.00571·cs.CL·December 2, 2020·6 cites

Denoising Pre-Training and Data Augmentation Strategies for Enhanced RDF Verbalization with Transformers

Sebastien Montella, Betty Fabre, Tanguy Urvoy, Johannes Heinecke, Lina, Rojas-Barahona

PDF

Open Access

TL;DR

This paper improves RDF triple verbalization by combining denoising pre-training and data augmentation strategies with Transformers, significantly enhancing text generation quality for both seen and unseen data categories.

Contribution

It introduces a novel approach that leverages augmented data and denoising pre-training to improve RDF-to-text verbalization with Transformer models.

Findings

01

BLEU score increases by up to 126.05% for unseen entities

02

Significant improvement in verbalization quality across categories

03

Demonstrates effectiveness of data augmentation in RDF verbalization

Abstract

The task of verbalization of RDF triples has known a growth in popularity due to the rising ubiquity of Knowledge Bases (KBs). The formalism of RDF triples is a simple and efficient way to store facts at a large scale. However, its abstract representation makes it difficult for humans to interpret. For this purpose, the WebNLG challenge aims at promoting automated RDF-to-text generation. We propose to leverage pre-trainings from augmented data with the Transformer model using a data augmentation strategy. Our experiment results show a minimum relative increases of 3.73%, 126.05% and 88.16% in BLEU score for seen categories, unseen entities and unseen categories respectively over the standard training.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Semantic Web and Ontologies

MethodsLinear Layer · Absolute Position Encodings · Position-Wise Feed-Forward Layer · Multi-Head Attention · Dense Connections · Attention Is All You Need · Adam · Softmax · Byte Pair Encoding · Label Smoothing