Making Robust Generalizers Less Rigid with Loss Concentration

Matthew J. Holland; Toma Hamada

arXiv:2408.03619·cs.LG·December 29, 2025

Making Robust Generalizers Less Rigid with Loss Concentration

Matthew J. Holland, Toma Hamada

PDF

1 Repo

TL;DR

This paper introduces a new training criterion that penalizes poor loss concentration to improve robustness in models, especially where the difficulty gap between data points is significant, surpassing traditional sharpness-aware methods.

Contribution

It proposes a flexible loss concentration penalty that can be combined with various loss transformations to enhance robustness beyond overparameterized neural networks.

Findings

01

Loss concentration penalty improves robustness in simpler models.

02

Combining the criterion with loss transformations like CVaR enhances tail performance.

03

Traditional sharpness-aware methods fail under models with large difficulty gaps.

Abstract

While the traditional formulation of machine learning tasks is in terms of performance on average, in practice we are often interested in how well a trained model performs on rare or difficult data points at test time. To achieve more robust and balanced generalization, methods applying sharpness-aware minimization to a subset of worst-case examples have proven successful for image classification tasks, but only using overparameterized neural networks under which the relative difference between "easy" and "hard" data points becomes negligible. In this work, we show how such a strategy can dramatically break down under simpler models where the difficulty gap becomes more extreme. As a more flexible alternative, instead of typical sharpness, we propose and evaluate a training criterion which penalizes poor loss concentration, which can be easily combined with loss transformations such…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

feedbackward/addro
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsSharpness-Aware Minimization