AD-DROP: Attribution-Driven Dropout for Robust Language Model   Fine-Tuning

Tao Yang; Jinghao Deng; Xiaojun Quan; Qifan Wang; Shaoliang Nie

arXiv:2210.05883·cs.CL·October 13, 2022

AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-Tuning

Tao Yang, Jinghao Deng, Xiaojun Quan, Qifan Wang, Shaoliang Nie

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces AD-DROP, a novel dropout method that selectively drops high-attribution attention positions during language model fine-tuning to improve generalization and reduce overfitting.

Contribution

The paper proposes Attribution-Driven Dropout (AD-DROP), a new regularization technique that targets high-attribution attention positions to enhance fine-tuning robustness.

Findings

01

AD-DROP improves performance across multiple benchmarks.

02

It acts as an effective regularizer against overfitting.

03

The method encourages reliance on low-attribution positions for predictions.

Abstract

Fine-tuning large pre-trained language models on downstream tasks is apt to suffer from overfitting when limited training data is available. While dropout proves to be an effective antidote by randomly dropping a proportion of units, existing research has not examined its effect on the self-attention mechanism. In this paper, we investigate this problem through self-attention attribution and find that dropping attention positions with low attribution scores can accelerate training and increase the risk of overfitting. Motivated by this observation, we propose Attribution-Driven Dropout (AD-DROP), which randomly discards some high-attribution positions to encourage the model to make predictions by relying more on low-attribution positions to reduce overfitting. We also develop a cross-tuning strategy to alternate fine-tuning and AD-DROP to avoid dropping high-attribution positions…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

taoyang225/ad-drop
jaxOfficial

Videos

AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-Tuning· slideslive

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Speech Recognition and Synthesis

MethodsDropout