Randomly Initialized Networks Can Learn from Peer-to-Peer Consensus

Esteban Rodr\'iguez-Betancourt; Edgar Casasola-Murillo

arXiv:2604.18390·cs.LG·May 1, 2026

Randomly Initialized Networks Can Learn from Peer-to-Peer Consensus

Esteban Rodr\'iguez-Betancourt, Edgar Casasola-Murillo

PDF

TL;DR

This paper demonstrates that even minimal, randomly initialized networks can improve their representations through peer-to-peer self-distillation, without complex mechanisms or pretext tasks.

Contribution

It isolates the effect of self-distillation in learning dynamics by training simple, randomly initialized networks, revealing their capacity to learn useful representations.

Findings

01

Randomly initialized networks can learn improved representations via self-distillation.

02

The effect varies with different hyperparameters.

03

Minimal setups can outperform random baselines on downstream tasks.

Abstract

In self-supervised learning, self-distilled methods have shown impressive performance, learning representations useful for downstream tasks and even displaying emergent properties. However, state-of-the-art methods usually rely on ensembles of complex mechanisms, with many design choices that are empirically motivated and not well understood. In this work, we explore the role of self-distillation within learning dynamics. Specifically, we isolate the effect of self-distillation by training a group of randomly initialized networks, removing all other common components such as projectors, predictors, and even pretext tasks. Our findings show that even this minimal setup can lead to learned representations with non-trivial improvements over a random baseline on downstream tasks. We also demonstrate how this effect varies with different hyperparameters and present a short analysis of what…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.