Training Neural Networks with Stochastic Hessian-Free Optimization

Ryan Kiros

arXiv:1301.3641·cs.LG·May 2, 2013·ICLR·27 cites

Training Neural Networks with Stochastic Hessian-Free Optimization

Ryan Kiros

PDF

Open Access

TL;DR

This paper introduces stochastic Hessian-free optimization, combining the benefits of Hessian-free methods and stochastic mini-batch training, enhanced with dropout to prevent overfitting, and demonstrates competitive results on classification and autoencoders.

Contribution

The paper adapts Hessian-free optimization to stochastic mini-batches and incorporates dropout, offering a scalable and effective training method for deep neural networks.

Findings

01

Achieves competitive performance on classification tasks.

02

Effective in training deep autoencoders.

03

Balances between SGD and full Hessian-free methods.

Abstract

Hessian-free (HF) optimization has been successfully used for training deep autoencoders and recurrent networks. HF uses the conjugate gradient algorithm to construct update directions through curvature-vector products that can be computed on the same order of time as gradients. In this paper we exploit this property and study stochastic HF with gradient and curvature mini-batches independent of the dataset size. We modify Martens' HF for these settings and integrate dropout, a method for preventing co-adaptation of feature detectors, to guard against overfitting. Stochastic Hessian-free optimization gives an intermediary between SGD and HF that achieves competitive performance on both classification and deep autoencoder experiments.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Advanced Neural Network Applications · Domain Adaptation and Few-Shot Learning

MethodsSolana Customer Service Number +1-833-534-1729 · Stochastic Gradient Descent