Generalization algorithm of multimodal pre-training model based on   graph-text self-supervised training

Zhangxiaobing; Tangzhenhao; Longzi; Fuxianghua

arXiv:2302.10315·cs.CL·February 22, 2023

Generalization algorithm of multimodal pre-training model based on graph-text self-supervised training

Zhangxiaobing, Tangzhenhao, Longzi, Fuxianghua

PDF

Open Access

TL;DR

This paper proposes a self-supervised multimodal pre-training algorithm that enhances neural machine translation by effectively utilizing visual information without requiring manual image annotation, leading to improved translation performance.

Contribution

It introduces a novel self-supervised training method that leverages web-sourced images to improve multimodal neural machine translation, overcoming data scarcity issues.

Findings

01

BLEU score increased by 0.5 on the global voice dataset

02

Effective use of web-sourced images improves translation accuracy

03

Overcomes visual data limitations in multimodal NMT

Abstract

Recently, a large number of studies have shown that the introduction of visual information can effectively improve the effect of neural machine translation (NMT). Its effectiveness largely depends on the availability of a large number of bilingual parallel sentence pairs and manual image annotation. The lack of images and the effectiveness of images have been difficult to solve. In this paper, a multimodal pre-training generalization algorithm for self-supervised training is proposed, which overcomes the lack of visual information and inaccuracy, and thus extends the applicability of images on NMT. Specifically, we will search for many pictures from the existing sentences through the search engine, and then through the relationship between visual information and text, do the self-supervised training task of graphics and text to obtain more effective visual information for text. We show…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques