LLaVA-SLT: Visual Language Tuning for Sign Language Translation

Han Liang; Chengyu Huang; Yuecheng Xu; Cheng Tang; Weicai Ye; Juze; Zhang; Xin Chen; Jingyi Yu; Lan Xu

arXiv:2412.16524·cs.CV·December 24, 2024·3 cites

LLaVA-SLT: Visual Language Tuning for Sign Language Translation

Han Liang, Chengyu Huang, Yuecheng Xu, Cheng Tang, Weicai Ye, Juze, Zhang, Xin Chen, Jingyi Yu, Lan Xu

PDF

Open Access

TL;DR

LLaVA-SLT introduces a multimodal model that leverages large language models and visual language embeddings to improve sign language translation accuracy without relying on costly gloss annotations.

Contribution

It presents a novel training framework combining linguistic pretraining, visual contrastive learning, and visual language tuning for sign language translation.

Findings

01

Outperforms state-of-the-art methods in SLT accuracy.

02

Closes the gap between gloss-free and gloss-based approaches.

03

Effective use of annotation-free data enhances performance.

Abstract

In the realm of Sign Language Translation (SLT), reliance on costly gloss-annotated datasets has posed a significant barrier. Recent advancements in gloss-free SLT methods have shown promise, yet they often largely lag behind gloss-based approaches in terms of translation accuracy. To narrow this performance gap, we introduce LLaVA-SLT, a pioneering Large Multimodal Model (LMM) framework designed to leverage the power of Large Language Models (LLMs) through effectively learned visual language embeddings. Our model is trained through a trilogy. First, we propose linguistic continued pretraining. We scale up the LLM and adapt it to the sign language domain using an extensive corpus dataset, effectively enhancing its textual linguistic knowledge about sign language. Then, we adopt visual contrastive pretraining to align the visual encoder with a large-scale pretrained text encoder. We…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHand Gesture Recognition Systems · Hearing Impairment and Communication · Subtitles and Audiovisual Media

MethodsADaptive gradient method with the OPTimal convergence rate · ALIGN