Improving Computational Complexity in Statistical Models with   Second-Order Information

Tongzheng Ren; Jiacheng Zhuo; Sujay Sanghavi; Nhat Ho

arXiv:2202.04219·stat.ML·April 15, 2022

Improving Computational Complexity in Statistical Models with Second-Order Information

Tongzheng Ren, Jiacheng Zhuo, Sujay Sanghavi, Nhat Ho

PDF

Open Access

TL;DR

This paper introduces a normalized gradient descent algorithm leveraging second-order information, significantly reducing the computational complexity for parameter estimation in singular statistical models, especially when the population loss is homogeneous.

Contribution

It proposes the NormGD algorithm that achieves logarithmic iteration complexity in sample size for homogeneous models, improving over traditional fixed step-size gradient descent.

Findings

01

NormGD reaches the statistical radius in logarithmic iterations for homogeneous models.

02

The algorithm achieves the optimal $\

03

contribution

Abstract

It is known that when the statistical models are singular, i.e., the Fisher information matrix at the true parameter is degenerate, the fixed step-size gradient descent algorithm takes polynomial number of steps in terms of the sample size $n$ to converge to a final statistical radius around the true parameter, which can be unsatisfactory for the application. To further improve that computational complexity, we consider the utilization of the second-order information in the design of optimization algorithms. Specifically, we study the normalized gradient descent (NormGD) algorithm for solving parameter estimation in parametric statistical models, which is a variant of gradient descent algorithm whose step size is scaled by the maximum eigenvalue of the Hessian matrix of the empirical loss function of statistical models. When the population loss function, i.e., the limit of the empirical…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Neural Networks and Applications · Markov Chains and Monte Carlo Methods