An Augmentation-Aware Theory for Self-Supervised Contrastive Learning

Jingyi Cui; Hongwei Wen; Yisen Wang

arXiv:2505.22196·cs.LG·May 29, 2025

An Augmentation-Aware Theory for Self-Supervised Contrastive Learning

Jingyi Cui, Hongwei Wen, Yisen Wang

PDF

Open Access 1 Video

TL;DR

This paper introduces an augmentation-aware theoretical framework for self-supervised contrastive learning, revealing how data augmentation influences learning performance and providing insights verified through experiments.

Contribution

It proposes the first augmentation-aware error bound for contrastive learning, explicitly linking data augmentation types to the learning risk.

Findings

01

Data augmentation significantly impacts the error bound in contrastive learning.

02

Certain augmentation methods can tighten or loosen the error bound.

03

Experimental results validate the theoretical predictions.

Abstract

Self-supervised contrastive learning has emerged as a powerful tool in machine learning and computer vision to learn meaningful representations from unlabeled data. Meanwhile, its empirical success has encouraged many theoretical studies to reveal the learning mechanisms. However, in the existing theoretical research, the role of data augmentation is still under-exploited, especially the effects of specific augmentation types. To fill in the blank, we for the first time propose an augmentation-aware error bound for self-supervised contrastive learning, showing that the supervised risk is bounded not only by the unsupervised risk, but also explicitly by a trade-off induced by data augmentation. Then, under a novel semantic label assumption, we discuss how certain augmentation methods affect the error bound. Lastly, we conduct both pixel- and representation-level experiments to verify our…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

An Augmentation-Aware Theory for Self-Supervised Contrastive Learning· slideslive

Taxonomy

TopicsDomain Adaptation and Few-Shot Learning · Face and Expression Recognition · Face recognition and analysis