Orthogonal Gradient Descent for Continual Learning

Mehrdad Farajtabar; Navid Azizan; Alex Mott; Ang Li

arXiv:1910.07104·cs.LG·October 17, 2019·39 cites

Orthogonal Gradient Descent for Continual Learning

Mehrdad Farajtabar, Navid Azizan, Alex Mott, Ang Li

PDF

Open Access

TL;DR

This paper introduces Orthogonal Gradient Descent (OGD), a method that mitigates catastrophic forgetting in continual learning by projecting gradients to preserve previous task performance without storing past data.

Contribution

The paper proposes a novel gradient projection technique that prevents forgetting in neural networks during continual learning without data rehearsal.

Findings

01

OGD effectively reduces catastrophic forgetting on benchmark tasks.

02

The method does not require storing previous data, enhancing privacy.

03

OGD maintains high performance across multiple sequential tasks.

Abstract

Neural networks are achieving state of the art and sometimes super-human performance on learning tasks across a variety of domains. Whenever these problems require learning in a continual or sequential manner, however, neural networks suffer from the problem of catastrophic forgetting; they forget how to solve previous tasks after being trained on a new task, despite having the essential capacity to solve both tasks if they were trained on both simultaneously. In this paper, we propose to address this issue from a parameter space perspective and study an approach to restrict the direction of the gradient updates to avoid forgetting previously-learned data. We present the Orthogonal Gradient Descent (OGD) method, which accomplishes this goal by projecting the gradients from new tasks onto a subspace in which the neural network output on previous task does not change and the projected…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDomain Adaptation and Few-Shot Learning · Multimodal Machine Learning Applications · COVID-19 diagnosis using AI