Improving Energy Conserving Descent for Machine Learning: Theory and   Practice

G. Bruno De Luca; Alice Gatti; Eva Silverstein

arXiv:2306.00352·cs.LG·June 2, 2023·1 cites

Improving Energy Conserving Descent for Machine Learning: Theory and Practice

G. Bruno De Luca, Alice Gatti, Eva Silverstein

PDF

Open Access 1 Repo

TL;DR

This paper introduces ECDSep, a novel energy-conserving optimization algorithm inspired by physical dynamical systems, demonstrating competitive performance in machine learning tasks and offering theoretical insights into its dynamics.

Contribution

The paper develops the theory of Energy Conserving Descent and introduces ECDSep, a new gradient-based optimizer that improves performance and simplifies hyper-parameter tuning for diverse problems.

Findings

01

ECDSep achieves competitive or superior results compared to SGD, Adam, and AdamW.

02

The method offers better control over the optimization dynamics.

03

Limitations suggest avenues for further enhancement.

Abstract

We develop the theory of Energy Conserving Descent (ECD) and introduce ECDSep, a gradient-based optimization algorithm able to tackle convex and non-convex optimization problems. The method is based on the novel ECD framework of optimization as physical evolution of a suitable chaotic energy-conserving dynamical system, enabling analytic control of the distribution of results - dominated at low loss - even for generic high-dimensional problems with no symmetries. Compared to previous realizations of this idea, we exploit the theoretical control to improve both the dynamics and chaos-inducing elements, enhancing performance while simplifying the hyper-parameter tuning of the optimization algorithm targeted to different classes of problems. We empirically compare with popular optimization methods such as SGD, Adam and AdamW on a wide range of machine learning problems, finding competitive…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

gbdl/ecdsep
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Adversarial Robustness in Machine Learning · Model Reduction and Neural Networks

MethodsAdam · Stochastic Gradient Descent · AdamW