ALBERT: A Lite BERT for Self-supervised Learning of Language   Representations

Zhenzhong Lan; Mingda Chen; Sebastian Goodman; Kevin Gimpel; Piyush; Sharma; Radu Soricut

arXiv:1909.11942·cs.CL·February 11, 2020·984 cites

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush, Sharma, Radu Soricut

PDF

Open Access 5 Repos 10 Models 1 Datasets

TL;DR

ALBERT introduces parameter-reduction techniques to create a more efficient BERT variant, achieving state-of-the-art results with fewer parameters and faster training, especially benefiting multi-sentence tasks.

Contribution

The paper proposes two novel parameter-reduction methods for BERT, improving scalability and training efficiency while maintaining or enhancing performance.

Findings

01

Achieves new state-of-the-art on GLUE, RACE, and SQuAD benchmarks.

02

Models are significantly smaller and faster to train than BERT-large.

03

Self-supervised loss focusing on inter-sentence coherence improves downstream task performance.

Abstract

Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer training times. To address these problems, we present two parameter-reduction techniques to lower memory consumption and increase the training speed of BERT. Comprehensive empirical evidence shows that our proposed methods lead to models that scale much better compared to the original BERT. We also use a self-supervised loss that focuses on modeling inter-sentence coherence, and show it consistently helps downstream tasks with multi-sentence inputs. As a result, our best model establishes new state-of-the-art results on the GLUE, RACE, and \squad benchmarks while having fewer parameters compared to BERT-large. The code and the pretrained models are…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Models

Datasets

dk-crazydiv/huggingface-modelhub
dataset· 32 dl
32 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Multimodal Machine Learning Applications

MethodsLinear Layer · SPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings · Weight Decay · Dropout · Attention Dropout · Linear Warmup With Linear Decay · BERT · Residual Connection · Adam · LAMB