Stochastic Gradient Hamiltonian Monte Carlo for Non-Convex Learning

Huy N. Chau; Miklos Rasonyi

arXiv:1903.10328·stat.ML·February 26, 2020

Stochastic Gradient Hamiltonian Monte Carlo for Non-Convex Learning

Huy N. Chau, Miklos Rasonyi

PDF

Open Access

TL;DR

This paper provides a non-asymptotic convergence analysis of Stochastic Gradient Hamiltonian Monte Carlo (SGHMC) for non-convex optimization, demonstrating its effectiveness with subsampling techniques in finding global minima.

Contribution

It offers the first non-asymptotic convergence analysis of SGHMC in non-convex settings, improving upon prior theoretical results.

Findings

01

Enhanced convergence guarantees for SGHMC in non-convex optimization

02

Improved theoretical bounds over previous analyses

03

Validation of SGHMC's effectiveness with subsampling techniques

Abstract

Stochastic Gradient Hamiltonian Monte Carlo (SGHMC) is a momentum version of stochastic gradient descent with properly injected Gaussian noise to find a global minimum. In this paper, non-asymptotic convergence analysis of SGHMC is given in the context of non-convex optimization, where subsampling techniques are used over an i.i.d dataset for gradient updates. Our results complement those of [RRT17] and improve on those of [GGZ18].

Equations358

F^{*} := x \in R^{d} min F (x), \mbox w h er e F (x) := E [f (x, Z)] = \int_{Z} f (x, z) μ (d z), x \in R^{d}

F^{*} := x \in R^{d} min F (x), \mbox w h er e F (x) := E [f (x, Z)] = \int_{Z} f (x, z) μ (d z), x \in R^{d}

E [F (X^{†})] - F^{*}

E [F (X^{†})] - F^{*}

x \in R^{d} min F_{z} (x), where F_{z} (x) := \frac{1}{n} i = 1 \sum n f (x, z_{i})

x \in R^{d} min F_{z} (x), where F_{z} (x) := \frac{1}{n} i = 1 \sum n f (x, z_{i})

d X_{t} = - \nabla F_{z} (X_{t}) d t + 2 β^{- 1} d B_{t},

d X_{t} = - \nabla F_{z} (X_{t}) d t + 2 β^{- 1} d B_{t},

d V_{t}

d V_{t}

d X_{t}

π_{z} (d x, d v) = \frac{1}{Γ _{z}} exp (- β (\frac{1}{2} ∥ v ∥^{2} + F_{z} (x))) d x d v

π_{z} (d x, d v) = \frac{1}{Γ _{z}} exp (- β (\frac{1}{2} ∥ v ∥^{2} + F_{z} (x))) d x d v

Γ_{z} = (\frac{2 π}{β})^{d /2} \int_{R^{d}} e^{- β F_{z} (x)} d x .

Γ_{z} = (\frac{2 π}{β})^{d /2} \int_{R^{d}} e^{- β F_{z} (x)} d x .

\overline{V}_{k + 1}^{λ}

\overline{V}_{k + 1}^{λ}

\overline{X}_{k + 1}^{λ}

E [g (x, U_{z})] = \nabla F_{z} (x), \forall x \in R^{d},

E [g (x, U_{z})] = \nabla F_{z} (x), \forall x \in R^{d},

V_{k + 1}^{λ}

V_{k + 1}^{λ}

X_{k + 1}^{λ}

E [F (X^{†})] - F^{*}

E [F (X^{†})] - F^{*}

W_{p} (μ, ν) = (π \in Π (μ, ν) in f \int_{R^{l}} ∥ x - y ∥^{p} d π (x, y))^{1/ p},

W_{p} (μ, ν) = (π \in Π (μ, ν) in f \int_{R^{l}} ∥ x - y ∥^{p} d π (x, y))^{1/ p},

∥ f (0, z) ∥ \leq A_{0}, ∥\nabla f (0, z) ∥ \leq B .

∥ f (0, z) ∥ \leq A_{0}, ∥\nabla f (0, z) ∥ \leq B .

∥\nabla f (x_{1}, z) - \nabla f (x_{2}, z) ∥ \leq M ∥ x_{1} - x_{2} ∥, \forall x_{1}, x_{2} \in R^{d} .

∥\nabla f (x_{1}, z) - \nabla f (x_{2}, z) ∥ \leq M ∥ x_{1} - x_{2} ∥, \forall x_{1}, x_{2} \in R^{d} .

⟨ x, f (x, z) ⟩ \geq m ∥ x ∥^{2} - b, \forall x \in R^{d}, z \in Z .

⟨ x, f (x, z) ⟩ \geq m ∥ x ∥^{2} - b, \forall x \in R^{d}, z \in Z .

∥ g (x_{1}, u) - g (x_{2}, u) ∥ \leq M ∥ x_{1} - x_{2} ∥, \forall x_{1}, x_{2} \in R^{d} .

∥ g (x_{1}, u) - g (x_{2}, u) ∥ \leq M ∥ x_{1} - x_{2} ∥, \forall x_{1}, x_{2} \in R^{d} .

E ∥ g (x, U_{z}) - \nabla F_{z} (x) ∥^{2} \leq 2 δ (M^{2} ∥ x ∥^{2} + B^{2}) .

E ∥ g (x, U_{z}) - \nabla F_{z} (x) ∥^{2} \leq 2 δ (M^{2} ∥ x ∥^{2} + B^{2}) .

\int_{R^{2 d}} e^{V (x, v)} d μ_{0} (x, v) < \infty,

\int_{R^{2 d}} e^{V (x, v)} d μ_{0} (x, v) < \infty,

g (x, U_{z}) = \frac{1}{ℓ} j = 1 \sum ℓ \nabla f (x, z_{I_{j}}),

g (x, U_{z}) = \frac{1}{ℓ} j = 1 \sum ℓ \nabla f (x, z_{I_{j}}),

d V (t, s, (v, x))

d V (t, s, (v, x))

d X (t, s, (v, x))

W_{p} ((V_{k}^{λ}, X_{k}^{λ}), (V (k, 0, (v_{0}, x_{0})), X (k, 0, (v_{0}, x_{0})))) \leq \tilde{C} (λ^{1/ (2 p)} + δ^{1/ (2 p)}) .

W_{p} ((V_{k}^{λ}, X_{k}^{λ}), (V (k, 0, (v_{0}, x_{0})), X (k, 0, (v_{0}, x_{0})))) \leq \tilde{C} (λ^{1/ (2 p)} + δ^{1/ (2 p)}) .

E [F (X_{k}^{λ})] - F^{*} \leq B_{1} + B_{2} + B_{3},

E [F (X_{k}^{λ})] - F^{*} \leq B_{1} + B_{2} + B_{3},

B_{1}

B_{1}

B_{2}

B_{3}

W_{p} (L (X_{k}), π_{z}) \leq ε

W_{p} (L (X_{k}), π_{z}) \leq ε

(λ^{1/ (2 p)} + δ^{1/ (2 p)}) \leq \frac{1}{2 C ~} ε, k \geq \frac{( 2 C ~ ) ^{2 p}}{c _{*}} \frac{1}{ε ^{2 p}} lo g (\frac{C _{*} ( W _{ρ} ( μ _{0} , π _{z} ) ) ^{1/ p}}{ε}) .

(λ^{1/ (2 p)} + δ^{1/ (2 p)}) \leq \frac{1}{2 C ~} ε, k \geq \frac{( 2 C ~ ) ^{2 p}}{c _{*}} \frac{1}{ε ^{2 p}} lo g (\frac{C _{*} ( W _{ρ} ( μ _{0} , π _{z} ) ) ^{1/ p}}{ε}) .

\tilde{C} (λ^{1/ (2 p)} + δ^{1/ (2 p)}) + C_{*} (W_{ρ} (μ_{0}, π_{z}))^{1/ p} \leq ε .

\tilde{C} (λ^{1/ (2 p)} + δ^{1/ (2 p)}) + C_{*} (W_{ρ} (μ_{0}, π_{z}))^{1/ p} \leq ε .

C_{*} (W_{ρ} (μ_{0}, π_{z}))^{1/ p} exp (- c_{*} k λ) \leq ε /2

C_{*} (W_{ρ} (μ_{0}, π_{z}))^{1/ p} exp (- c_{*} k λ) \leq ε /2

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Markov Chains and Monte Carlo Methods · Medical Imaging Techniques and Applications

Full text

Stochastic Gradient Hamiltonian Monte Carlo for Non-Convex Learning ††thanks: Both authors were supported by the NKFIH (National Research, Development and Innovation Office, Hungary) grant KH 126505 and the “Lendület” grant LP 2015-6 of the Hungarian Academy of Sciences. The authors thank Minh-Ngoc Tran for helpful discussions.

Huy N. Chau

Miklós Rásonyi

Abstract

Stochastic Gradient Hamiltonian Monte Carlo (SGHMC) is a momentum version of stochastic gradient descent with properly injected Gaussian noise to find a global minimum. In this paper, non-asymptotic convergence analysis of SGHMC is given in the context of non-convex optimization, where subsampling techniques are used over an i.i.d dataset for gradient updates. Our results complement those of [RRT17] and improve on those of [GGZ18].

1 Introduction

Let $(\Omega,\mathcal{F},P)$ be a probability space where all the random objects of this paper will be defined. The expectation of a random variable $X$ with values in a Euclidean space will be denoted by $E[X]$ .

We consider the following optimization problem

[TABLE]

and $Z$ is a random element in some measurable space $\mathcal{Z}$ with an unknown probability law $\mu$ . The function $x\mapsto f(x,z)$ is assumed continuously differentiable (for each $z$ ) but it can possibly be non-convex. Suppose that one has access to i.i.d samples $\mathbf{Z}=(Z_{1},...,Z_{n})$ drawn from $\mu$ , where $n\in\mathbb{N}$ is fixed. Our goal is to compute an approximate minimizer $X^{\dagger}$ such that the population risk

[TABLE]

is minimized, where the expectation is taken with respect to the training data $\mathbf{Z}$ and additional randomness generating $X^{\dagger}$ .

Since the distribution of $Z_{i},i\in\mathbb{N}$ is unknown, we consider the empirical risk minimization problem

[TABLE]

using the dataset $\mathbf{z}:=\{z_{1},...,z_{n}\}$

Stochastic gradient algorithms based on Langevin Monte Carlo have gained more attention in recent years. Two popular algorithms are Stochastic Gradient Langevin Dynamics (SGLD) and Stochastic Gradient Hamiltonian Monte Carlo (SGHMC). First, we summarize the use of SGLD in optimization, as presented in [RRT17]. Consider the overdamped Langevin stochastic differential equation

[TABLE]

where $(B_{t})_{t\geq 0}$ is the standard Brownian motion in $\mathbb{R}^{d}$ and $\beta>0$ is the inverse temperature parameter. Under suitable assumptions on $f$ , the SDE (3) admits the Gibbs measure $\pi_{\mathbf{z}}(dx)\propto\exp(-\beta F_{\mathbf{z}}(x))$ as its unique invariant distribution. In addition, it is known that for sufficiently big $\beta$ , the Gibbs distribution concentrates around global minimizers of $F_{\mathbf{z}}$ . Therefore, one can use the value of $X_{t}$ from (3), (or from its discretized counterpart SGLD), as an approximate solution to the empirical risk problem, provided that $t$ is large and temperature is low.

In this paper, we consider the underdamped (second-order) Langevin diffusion

[TABLE]

where $(X_{t})_{t\geq 0},(V_{t})_{t\geq 0}$ model the position and the momentum of a particle moving in a field of force $F_{\mathbf{z}}$ with random force given by Gaussian noise. It is shown that under some suitable conditions for $F_{\mathbf{z}}$ , the Markov process $(X,V)$ is ergodic and has a unique stationary distribution

[TABLE]

where $\Gamma_{\mathbf{z}}$ is the normalizing constant

[TABLE]

It is easy to observe that the $x$ -marginal distribution of $\pi_{\mathbf{z}}(dx,dv)$ is the invariant distribution $\pi_{\mathbf{z}}(dx)$ of (3). We consider the first order Euler discretization of (4), (5), also called Stochastic Gradient Hamiltonian Monte Carlo (SGHMC), given as follows

[TABLE]

where $\lambda>0$ is a step size parameter and $(\xi_{k})_{k\in\mathbb{N}}$ is a sequence of i.i.d standard Gaussian random vectors in $\mathbb{R}^{d}$ . The initial condition $v_{0},x_{0}$ may be random, but independent of $(\xi_{k})_{k\in\mathbb{N}}$ .

In certain contexts, the full knowledge of the gradient $F_{\mathbf{z}}$ is not available, however, using the dataset $\mathbf{z}$ , one can construct its unbiased estimates. In what follows, we adopt the general setting given by [RRT17]. Let $\mathcal{U}$ be a measurable space, and $g:\mathbb{R}^{d}\times\mathcal{U}\to\mathbb{R}^{d}$ such that for any $\mathbf{z}\in\mathcal{Z}^{n}$ ,

[TABLE]

where $U_{\mathbf{z}}$ is a random element in $\mathcal{U}$ with probability law $Q_{\mathbf{z}}$ . Conditionally on $\mathbf{Z}=\mathbf{z}$ , the SGHMC algorithm is defined by

[TABLE]

where $(U_{\mathbf{z},k})_{k\in\mathbb{N}}$ is a sequence of i.i.d. random elements in $\mathcal{U}$ with law $Q_{\mathbf{z}}$ . We also assume from now on that $v_{0},x_{0},(U_{\mathbf{z},k})_{k\in\mathbb{N}},(\xi_{k})_{k\in\mathbb{N}}$ are independent.

Our ultimate goal is to find approximate global minimizers to the problem (1). Let $X^{\dagger}:=X^{\lambda}_{k}$ be the output of the algorithm (9),(10) after $k\in\mathbb{N}$ iterations, and $(\widehat{X}^{*}_{\mathbf{z}},\widehat{V}^{*}_{\mathbf{z}})$ be such that $\mathcal{L}(\widehat{X}^{*}_{\mathbf{z}},\widehat{V}^{*}_{\mathbf{z}})=\pi_{\mathbf{z}}$ . The excess risk is decomposed as follows, see also [RRT17],

[TABLE]

The remaining part of the present paper is about finding bounds for these errors. Section 2 summarizes technical conditions and the main results. Comparison of our contributions to previous studies is discussed in Section 3. Proofs are given in Section 4.

Notation and conventions. For $l\geq 1$ , scalar product in $\mathbb{R}^{l}$ is denoted by $\langle\cdot,\cdot\rangle$ . We use $\|\cdot\|$ to denote the Euclidean norm (where the dimension of the space may vary). $\mathcal{B}(\mathbb{R}^{l})$ denotes the Borel $\sigma$ - field of $\mathbb{R}^{l}$ . For any $\mathbb{R}^{l}$ -valued random variable $X$ and for any $1\leq p<\infty$ , let us set $\|X\|_{p}:=E^{1/p}\|X\|^{p}$ . We denote by $L^{p}$ the set of $X$ with $\|X\|_{p}<\infty$ . The Wasserstein distance of order $p\in[1,\infty)$ between two probability measures $\mu$ and $\nu$ on $\mathcal{B}(\mathbb{R}^{l})$ is defined by

[TABLE]

where $\Pi(\mu,\nu)$ is the set of couplings of $(\mu,\nu)$ , see e.g. [Vil08]. For two $\mathbb{R}^{l}$ -valued random variables $X$ and $Y$ , we denote $\mathfrak{W}_{2}(X,Y):=\mathcal{W}_{2}(\mathcal{L}(X),\mathcal{L}(Y))$ , where $\mathcal{L}(X)$ is the law of $X$ . We do not indicate $l$ in the notation and it may vary.

2 Asumptions and main results

The following conditions are required throughout the paper.

Assumption 2.1.

The function $f$ is continuously differentiable, takes non-negative values, and there are constants $A_{0},B\geq 0$ such that for any $z\in\mathcal{Z}$ ,

[TABLE]

Assumption 2.2.

There is $M>0$ such that, for each $z\in\mathcal{Z}$ ,

[TABLE]

Assumption 2.3 (Dissipative).

There exist constants $m>0,b\geq 0$ such that

[TABLE]

Assumption 2.4.

For each $u\in\mathcal{U}$ , it holds that $\|g(0,u)\|\leq B$ and

[TABLE]

Assumption 2.5.

There exists a constant $\delta>0$ such that for every $\mathbf{z}\in\mathcal{Z}^{n}$ ,

[TABLE]

Assumption 2.6.

The law $\mu_{0}$ of the initial state $(x_{0},v_{0})$ satisfies

[TABLE]

where $\mathcal{V}$ is the Lyapunov function defined in (17) below.

Remark 2.7.

If the set of global minimizers is bounded, we can always redefine the function $f$ to be quadratic outside a compact set containing the origin while maintaining its minimizers. Hence, Assumption 2.3 can be satisfied in practice. Assumption 2.4 means that the estimated gradient is also Lipschitz when using the same training dataset. For example, at each iteration of SGHMC, we may sample uniformly with replacement a random minibatch of size $\ell$ . Then we can choose $U_{\mathbf{z}}=(z_{I_{1}},...,z_{I_{\ell}})$ where $I_{1},...,I_{\ell}$ are i.i.d random variables having distribution $\text{Uniform}(\{1,...,n\})$ . The gradient estimate is thus

[TABLE]

which is clearly unbiased and Assumption 2.4 will be satisfied whenever Assumptions 2.2 and 2.1 are in force. Assumption 2.5 controls the variance of the gradient estimate.

An auxiliary continuous time process is needed in the subsequent analysis. For a step size $\lambda>0$ , denote by $B^{\lambda}_{t}:=\frac{1}{\sqrt{\lambda}}B_{\lambda t}$ the scaled Brownian motion. Let $\widehat{V}(t,s,(v,x)),\widehat{X}(t,s,(v,x))$ be the solutions of

[TABLE]

with initial condition $\widehat{V}_{s}=v,\widehat{X}_{s}=x$ where $v,x$ may be random but independent of $(B^{\lambda}_{t})_{t\geq 0}$ .

Our first result tracks the discrepancy between the SGHMC algorithm (9), (10) and the auxiliary processes (13), (14).

Theorem 2.8.

Let $1\leq p\leq 2$ . There exists a constant $\tilde{C}>0$ such that for all $k\in\mathbb{N}$ ,

[TABLE]

Proof.

The proof of this theorem is given in Section 4.2. ∎

The following is the main result of the paper.

Theorem 2.9.

Let $1<p\leq 2$ . Suppose that the SGHMC iterates $(V^{\lambda}_{k},X^{\lambda}_{k})$ are defined by (9), (10). The expected population risk can be bounded as

[TABLE]

where

[TABLE]

where $\tilde{C},C_{*},c_{*},c_{LS}$ are appropriate constants and $\mathcal{W}_{\rho}$ is the metric defined in (20) below.

Proof.

The proof of this theorem is given in Section 4.3. ∎

Corollary 2.10.

Let $1\leq p\leq 2,\varepsilon>0$ We have

[TABLE]

whenever

[TABLE]

Proof.

From the proof of Theorem 2.9, or more precisely from (46), we need to choose $\lambda$ and $k$ such that

[TABLE]

First, we choose $\lambda$ and $\delta$ so that $\tilde{C}(\lambda^{1/(2p)}+\delta^{1/(2p)})<\varepsilon/2$ and then

[TABLE]

will hold for $k$ large enough. ∎

3 Related work and our contributions

Non-asymptotic convergence rate Langevin dynamics based algorithms for approximate sampling log-concave distributions are intensively studied in recent years. For example, overdamped Langevin dynamics are discussed in [WT11], [Dal17b], [DM16], [DK17], [DM17] and others. Recently, [BCM*+*18] treats the case of non-i.i.d. data streams with a certain mixing property. Underdamped Langevin dynamics are examined in [CFG14], [Nea11], [CCBJ17], etc. Further analysis on HMC are discussed on [BBLG17], [Bet17]. Subsampling methods are applied to speed up HMC for large datasets, see [DQK*+*17], [QKVT18].

The use of momentum to accelerate optimization methods are discussed intensively in literature, for example [AP16]. In particular, performance of SGHMC is experimentally proved better than SGLD in many applications, see [CDC15], [CFG14]. An important advantage of the underdamped SDE is that convergence to its stationary distribution is faster than that of the overdamped SDE in the $2$ -Wasserstein distance, as shown in [EGZ17].

Finding an approximate minimizer is similar to sampling distributions concentrate around the true minimizer. This well known connection gives rise to the study of simulated annealing algorithms, see [Hwa80], [Gid85], [Haj85], [CHS87], [HKS89], [GM91], [GM93]. Recently, there are many studies further investigate this connection by means of non asymptotic convergence of Langevin based algorithms and in stochastic non-convex optimization and large-scale data analysis, [CCG*+*16], [Dal17a].

Relaxing convexity is a more challenging issue. In [CCAY*+*18], the problem of sampling from a target distribution $\exp(-F(x))$ where $F$ is L-smooth everywhere and $m$ -strongly convex outside a ball of finite radius is considered. They provide upper bounds for the number of steps to be within a given precision level $\varepsilon$ of the 1-Wasserstein distance between the HMC algorithm and the equilibrium distribution. In a similar setting, [MMS18] obtains bounds in both the $\mathcal{W}_{1}$ and $\mathcal{W}_{2}$ distances for overdamped Langevin dynamics with stochastic gradients. [XCZG18] studies the convergence of the SGLD algorithm and the variance reduced SGLD to global minima of nonconvex functions satisfying the dissipativity condition.

Our work continues these lines of research, the most similar setting to ours is the recent paper [GGZ18]. We summarize our contributions below:

•

Diffusion approximation. In Lemma 10 of [GGZ18], the upper bound for the 2-Wasserstein distance between the SGHMC algorithm at step $k$ and underdamped SDE at time $t=k\lambda$ is (up to constants) given by

[TABLE]

which depends on the number of iteration $k$ . Therefore obtaining a precision $\varepsilon$ requires a careful choice of $k,\lambda$ and even $k\lambda$ . By introducing the auxiliary SDEs (13), (14), we are able to achieve the rate

[TABLE]

see Theorem 2.8 for the case $p=2$ . This upper bound is better in the number of iterations and hence, improves Lemma 10 of [GGZ18]. Our analysis for variance of the algorithm is also different. The iteration does not accumulate mean squared errors, as the number of step goes to infinity.

•

Our proof for Theorem 2.8 is relatively simple and we do not need to adopt the techniques of [RRT17] which involve heavy functional analysis, e.g. the weighted Csiszár - Kullback - Pinsker inequalities in [BV05] is not needed.

•

If we consider the $p$ -Wasserstein distance for $1<p\leq 2$ , in particular, when $p\to 1$ , Theorem 2.9 gives tighter bounds, compared to Theorem 2 of [GGZ18].

•

Dependence structure of the dataset in the sampling mechanism, can be arbitrary, see the proof of Theorem 2.8. The i.i.d assumption on dataset is used only for the generalization error. We could also incorporate non-i.i.d data in our analysis, see Remark 4.5, but this is left for further research.

4 Proofs

4.1 A contraction result

In this section, we recall a contraction result of [EGZ17]. First, it should be noticed that the constant $u$ and the function $U$ in their paper are $\beta^{-1}$ and $\beta F_{\mathbf{z}}$ in the present paper, respectively. Here, the subscript $c$ stands for “contraction”. Using the upper bound of Lemma 5.1 for $f$ below, there exist constants $\lambda_{c}\in\left(0,\min\{1/4,m/(M+2B+\gamma^{2}/2)\}\right)$ small enough and $A_{c}\geq\beta/2(b+2B+A_{0})$ such that

[TABLE]

Therefore, Assumption 2.1 of [EGZ17] is satisfied, noting that $L_{c}:=\beta M$ and

[TABLE]

We define the Lyapunov function

[TABLE]

For any $(x_{1},v_{1}),(x_{2},v_{2})\in\mathbb{R}^{2d}$ , we set

[TABLE]

where $\alpha_{c},\varepsilon_{c}>0$ are suitable positive constants to be fixed later and $h:[0,\infty)\to[0,\infty)$ is continuous, non-decreasing concave function such that $h(0)=0$ , $h$ is $C^{2}$ on $(0,R_{1})$ for some constant $R_{1}>0$ with right-sided derivative $h^{\prime}_{+}(0)=1$ and left-sided derivative $h^{\prime}_{-}(R_{1})>0$ and $h$ is constant on $[R_{1},\infty)$ . For any two probability measures $\mu,\nu$ on $\mathbb{R}^{2d}$ , we define

[TABLE]

Note that $\rho$ and $\mathcal{W}_{\rho}$ are semimetrics but not necessarily metrics. A result from [EGZ17] is recalled below.

For a probability measure $\mu$ on $\mathcal{B}(\mathbb{R}^{2d})$ , we denote by $\mu p_{t}$ the law of $(V_{t},X_{t})$ when $\mathcal{L}(V_{0},X_{0})=\mu$ .

Theorem 4.1.

There exists a continuous non-decreasing concave function $h$ with $h(0)=0$ such that for all probability measures $\mu,\nu$ on $\mathbb{R}^{2d}$ , and $1\leq p\leq 2$ , we have

[TABLE]

where the following relations hold:

[TABLE]

The function $h$ is constant on $[R_{1},\infty)$ , $C^{2}$ on $(0,R_{1})$ with

[TABLE]

and $\eta_{c}$ satisfies $\alpha_{c}=(1+\eta_{c})L_{c}\beta^{-1}\gamma^{-2}$ .

Proof.

From (5.15) of [EGZ17], we get

[TABLE]

Furthermore, from the proof of Corollary 2.6 of [EGZ17], if $r:=r((x_{1},v_{1}),(x_{2},v_{2}))\leq\min\{1,R_{1}\}$ ,

[TABLE]

and if $r\geq\min\{1,R_{1}\}$ then

[TABLE]

These bounds and Theorem 2.3 of [EGZ17] imply that

[TABLE]

The proof is complete. ∎

It should be emphasized that $(\widehat{V}(t,0,(v_{0},x_{0})),\widehat{X}(t,0,(v_{0},x_{0})))=(V_{\lambda t},X_{\lambda t})$ , and consequently, $(\widehat{V}(t,0,(v_{0},x_{0})),\widehat{X}(t,0,(v_{0},x_{0})))$ contracts at the rate $\exp(-c_{*}\lambda t)$ .

4.2 Proof of Theorem 2.8

Here, we summarize our approach. For a given step size $\lambda>0$ , we divide the time axis into intervals of length $T=\lfloor 1/\lambda\rfloor$ . For each time step $k\in[nT,(n+1)T],n\in\mathbb{N}$ , we compare the SGHMC to the version with exact gradients relying on the Doob inequality, and then compare the later to the auxiliary continuous-time diffusion $(\widehat{V}(k,0,(v_{0},x_{0})),\widehat{X}(k,0,(v_{0},x_{0})))$ with the scaled Brownian motion. At this stage we reply on the contraction result from [EGZ17] and uniform boundedness of the Langevin diffusion and its discrete time versions. Since the auxiliary dynamics evolves slower than the original Langevin dynamics, or more precisely at the same speed as that of the SGHCM, our upper bounds do not accumulate errors and are independent from the number of iterations.

Proof.

For each $k\in\mathbb{N}$ , we define

[TABLE]

Let $\tilde{v},\tilde{x}$ be $\mathbb{R}^{d}$ -valued random variables satisfying Assumption 2.6. For $0\leq i\leq j$ , we recursively define $\tilde{V}^{\lambda}(i,i,(\tilde{v},\tilde{x})):=\tilde{v}$ , $\tilde{X}^{\lambda}(i,i,(\tilde{v},\tilde{x})):=\tilde{x}$ and

[TABLE]

Let $T:=\lfloor 1/\lambda\rfloor$ . For each $n\in\mathbb{N}$ , and for each $nT\leq k<(n+1)T$ , we set

[TABLE]

For each $n\in\mathbb{N}$ , it holds by definition that $V^{\lambda}_{nT}=\tilde{V}^{\lambda}_{nT}$ and the triangle inequality implies for $nT\leq k<(n+1)T$ ,

[TABLE]

and

[TABLE]

Denote $g_{k,nT}(x):=E\left[g(x,U_{\mathbf{z},k})|\mathcal{H}_{nT}\right],x\in\mathbb{R}^{d}$ . By Assumption 2.4, the estimation continues as follows

[TABLE]

Using (25), one obtains

[TABLE]

noting that $T\lambda\leq 1.$ Therefore, the estimation in (26) continues as

[TABLE]

Applying the discrete-time version of Grönwall’s lemma and taking squares, noting also that $(x+y)^{2}\leq 2(x^{2}+y^{2}),x,y\in\mathbb{R}$ yield

[TABLE]

where

[TABLE]

Taking conditional expectation with respect to $\mathcal{H}_{nT}$ , the estimation becomes

[TABLE]

Since the random variables $U_{\mathbf{z},i}$ are independent, the sequence of random variables $g(\tilde{X}^{\lambda}_{i},U_{\mathbf{z},i})-g_{i,nT}(\tilde{X}^{\lambda}_{i})$ , $nT\leq i<(n+1)T$ are independent conditionally on $\mathcal{H}_{nT}$ , noting that $\tilde{X}^{\lambda}_{i}$ is measurable with respect to $\mathcal{H}_{nT}$ . In addition, they have zero mean by the tower property of conditional expectation. By Assumption 2.4,

[TABLE]

and thus

[TABLE]

by the independence of $U_{\mathbf{z},i},i>nT$ from $\mathcal{H}_{nT}$ . Doob’s inequality and (29) imply

[TABLE]

Taking one more expectation and using Lemma 5.3 give

[TABLE]

By Lemma 4.3, we have $E[\Xi^{2}_{n}]<2T^{2}\delta(M^{2}C^{a}_{x}+B^{2})$ , and therefore,

[TABLE]

where we define

[TABLE]

Consequently, we have from (25)

[TABLE]

Let $\tilde{V}^{int}$ and $\tilde{X}^{int}$ be the continuous-time interpolation of $\tilde{V}^{\lambda}_{k}$ , and of $\tilde{X}^{\lambda}_{k}$ on $[nT,(n+1)T)$ , respectively,

[TABLE]

with the initial conditions $\tilde{V}^{int}_{nT}=\tilde{V}_{nT}=V^{\lambda}_{nT}$ and $\tilde{X}^{int}_{nT}=\tilde{X}_{nT}=X^{\lambda}_{nT}$ . For each $n\in\mathbb{N}$ and for $nT\leq t<(n+1)T$ , define also

[TABLE]

where the dynamics of $\widehat{V},\widehat{X}$ are given in (13), (14). In this way, the processes $(\widehat{V}_{t})_{t\geq 0},(\widehat{X}_{t})_{t\geq 0}$ are right continuous with left limits. From Lemma 4.4, we obtain for $nT\leq t<(n+1)T$

[TABLE]

Combining (30), (31) and (35) gives

[TABLE]

Define $\widehat{A}_{t}=(\widehat{V}_{t},\widehat{X}_{t})$ and $\widehat{B}(t,s,(v_{s},x_{s}))=(\widehat{V}(t,s,(v_{s},x_{s})),\widehat{X}(t,s,(v_{s},x_{s})))$ for $s\leq t$ and $v_{s},x_{s}$ are $\mathbb{R}^{d}$ -valued random variables. The triangle inequality and Theorem 4.1 imply that for $nT\leq t<(n+1)T$ , and for $1\leq p\leq 2$ ,

[TABLE]

noting the rate of contraction of $(\widehat{V}_{t},\widehat{X}_{t})$ is $e^{-c_{*}\lambda t}$ . Using Lemma 5.4, we obtain

[TABLE]

where

[TABLE]

Now, we compute

[TABLE]

In $L^{2}$ norm, the first and second terms of (38) is bounded by $(c_{2}+c_{7})\sqrt{\lambda}+c_{3}\sqrt{\delta}$ , see (36) and the fifth term is estimated by $\sqrt{\lambda}$ . We consider the third term in (38). From the dynamics of $\widehat{V}$ , we find that for $iT-1\leq t\leq iT$ ,

[TABLE]

Hölder’s inequality yields

[TABLE]

where the last inequality uses Lemma 5.3 and Assumption 2.2 and $c_{14}:=3\gamma^{2}C^{c}_{v}+6M^{2}C^{c}_{x}+6B^{2}+6\gamma\beta^{-1}$ . For the fourth term of (38), we have

[TABLE]

where the last inequality uses Assumption 2.5, Lemma 5.3, and (36) and $c_{15}:=\max\{2(M^{2}C^{a}_{x}+B^{2})+4M^{2}c_{3}^{2},4M^{2}(c_{2}+c_{7})^{2}\}$ . A similar estimate holds for

[TABLE]

Letting $c_{16}:=\max\{(c_{2}+c_{7}),c_{3},\sqrt{c_{14}},\sqrt{c_{15}}\}$ , the estimation (37) continues as

[TABLE]

Therefore, from (30), (35), (39), the triangle inequality implies for $nT\leq k<(n+1)T$ ,

[TABLE]

where $\tilde{C}=2\max\{c_{2},c_{3},c_{7},C_{*}\left(c_{18}c_{16}\right)^{1/p}\frac{e^{-c_{*}}}{1-e^{-c_{*}}}\}$ . The proof is complete. ∎

Remark 4.2.

It is important to remark from the proof above that the data structure of $\mathbf{Z}$ can be arbitrary, and only the independence of random elements $U_{\mathbf{z},k},k\in\mathbb{N}$ is used.

Lemma 4.3.

The quantity $\Xi_{n}$ defined in (28) has second moments and

[TABLE]

Proof.

Noting that for each $nT\leq i<(n+1)T-1$ , the random variable $\tilde{X}^{\lambda}_{i}$ is $\mathcal{H}_{nT}$ -measurable. Using Assumption 2.5, the Cauchy–Schwarz inequality implies

[TABLE]

where the last inequality uses Lemma 5.3. ∎

This lemma provides variance control for the algorithm. Each term in $\Xi_{n}$ has an error of order $\delta$ , the total variance in $\Xi_{n}$ is of order $T\delta$ . However, unlike [RRT17], [GGZ18], our technique does not accumulate variance errors over time, as shown in (30). Recently in [BCM*+*18], the authors imposed no condition for variance of the estimated gradient, but employ the conditional $L$ -mixing property of data stream, and hence variance is controlled by the decay of mixing property, see their Lemma 8.6.

Lemma 4.4.

For every $nT\leq t<(n+1)T$ , it holds that

[TABLE]

Proof.

Noting that $\tilde{V}^{int}_{nT}=\widehat{V}_{nT}=V^{\lambda}_{nT}$ , we use the triangle inequality and Assumption 2.2 to estimate

[TABLE]

For notational convenience, we define for every $nT\leq t<(n+1)T$

[TABLE]

Then (40) becomes

[TABLE]

Furthermore,

[TABLE]

We estimate

[TABLE]

Noting that $0\leq t-\lfloor t\rfloor\leq 1$ , the Cauchy-Schwarz inequality and Lemma 5.1 imply

[TABLE]

Taking expectation both sides and noting that $(\tilde{V}^{int}_{k},\tilde{X}^{int}_{k})$ has the same distribution as $(\tilde{V}^{\lambda}_{k},\tilde{X}^{\lambda}_{k}),k\in\mathbb{N}$ , Lemma 5.3 leads to

[TABLE]

for $c_{8}:=3\gamma^{2}C^{a}_{v}+6M^{2}C^{a}_{x}+6B^{2}+6\gamma\beta^{-1}$ . Similarly,

[TABLE]

Taking squares and expectation of (41), (42), applying (43), (44) we obtain for $nT\leq t<(n+1)T$

[TABLE]

where $c_{9}:=\max\{4\gamma^{2}c_{8}+4M^{2}C^{a}_{v},2c_{8}\}$ . Summing up two inequalities yields

[TABLE]

where $c_{10}:=\max\{4\gamma^{2}+2,4M^{2}\}$ and then Gronwall’s lemma shows

[TABLE]

noting that $t\mapsto E[I^{2}_{t}+J^{2}_{t}]$ is continuous. The proof is complete by setting $c_{7}=\sqrt{2c_{9}e^{c_{10}}}$ , which is of order $\sqrt{d}$ . ∎

4.3 Proof of Theorem 2.9

Denote $\mu_{\mathbf{z},k}:=\mathcal{L}((V^{\lambda}_{k},X^{\lambda}_{k})|\mathbf{Z}=\mathbf{z})$ . Let $(\widehat{X},\widehat{V})$ and $(\widehat{X}^{*},\widehat{V}^{*})$ be such that $\mathcal{L}((\widehat{X},\widehat{V})|\mathbf{Z}=\mathbf{z})=\mu_{\mathbf{z},k}$ and $\mathcal{L}(\widehat{X}^{*}_{\mathbf{z}},\widehat{V}^{*}_{\mathbf{z}})=\pi_{\mathbf{z}}$ . We decompose the population risk by

[TABLE]

4.3.1 The first term $\mathcal{T}_{1}$

The first term in the right hand side of (45) is rewritten as

[TABLE]

where $\mu^{\otimes n}$ is the product of laws of independent random variables $Z_{1},...,Z_{n}$ . By Assumptions 2.1 and 2.2, the function $F_{\mathbf{z}}$ satisfies $\|\nabla F_{\mathbf{z}}(x)\|\leq M\|x\|+B.$ Using Lemma 5.2, we have

[TABLE]

where $p>1,q\in\mathbb{N},1/p+1/(2q)=1,$

[TABLE]

by Lemma 5.5. On the other hand, Theorems 2.8 and 4.1 imply

[TABLE]

Therefore, an upper bound for $\mathcal{T}_{1}$ is given by

[TABLE]

4.3.2 The second term $\mathcal{T}_{2}$

Since the $x$ -marginal of $\pi_{\mathbf{z}}(dx,dv)$ is $\pi_{\mathbf{z}}(dx)$ , the Gibbs measure of (3), we compute

[TABLE]

Therefore the argument in [RRT17] is adopted,

[TABLE]

The constant $c_{LS}$ comes from the logarithmic Sobolev inequality for $\pi_{\mathbf{z}}$ and

[TABLE]

where $\lambda_{*}$ is the uniform spectral gap for the overdamped Langevin dynamics

[TABLE]

Remark 4.5.

One can also find an upper bound for $\mathcal{T}_{2}$ when the data $\mathbf{z}$ is a realization of some non-Makovian processes. For example, if we assume that $f$ is Lipschitz on the second variable $z$ and $\mathbf{Z}$ satisfies a certain mixing property discussed in [CKRS16]) then the term $\mathcal{T}_{2}$ is bounded by $1/\sqrt{n}$ times a constant, see Theorem 2.5 therein.

4.3.3 The third term $\mathcal{T}_{3}$

For the third term, we follow [RRT17]. Let $x^{*}$ be any minimizer of $F(x)$ . We compute

[TABLE]

where the last inequality comes from Proposition 3.4 of [RRT17]. The condition $\beta\geq 2m$ is not used here, see the explanation in Lemma 16 of [GGZ18].

5 Technical lemmas

Lemma 5.1.

Under Assumptions 2.1, 2.2, for any $x\in\mathbb{R}^{d}$ and $z\in\mathcal{U}$ ,

[TABLE]

and

[TABLE]

Proof.

See Lemma 2 of [RRT17]. ∎

The next lemma generalizes continuity for functions of quadratic growth in Wasserstein distances given in [PW16].

Lemma 5.2.

Let $\mu,\nu$ be two probability measures on $\mathbb{R}^{2d}$ with finite second moments and let $G:\mathbb{R}^{2d}\to\mathbb{R}$ be a $C^{1}$ function with

[TABLE]

for some $c_{1}>0,c_{2}\geq 0$ . Then for $p>1,q>1$ such that $1/p+1/q=1$ , we have

[TABLE]

where

[TABLE]

Proof.

Using the Cauchy-Schwartz inequality, we compute

[TABLE]

Then for any $\xi\in\Pi(\mu,\nu)$ we have

[TABLE]

Since this inequality holds true for any $\xi\in\Pi(\mu,\nu)$ , the proof is complete. ∎

Lemma 5.3.

The continuous time processes (4),(5) are uniformly bounded in $L^{2}$ , more precisely,

[TABLE]

For $0<\lambda\leq\min\left\{\frac{\gamma}{K_{2}}\left(\frac{d+A_{c}}{\beta},\frac{\gamma\lambda_{c}}{2K_{1}}\right)\right\}$ , where

[TABLE]

and

[TABLE]

the SGHMC (9),(10) satisfy

[TABLE]

Furthermore, the processes defined in (24), (34) are also uniformly bounded in $L^{2}$ with the upper bounds $C^{c}_{v},C^{c}_{x},C^{a}_{v},C^{a}_{x}$ , respectively.

Proof.

The uniform boundedness in $L^{2}$ of the processes in (4), (5), (9), (10) are given in Lemma 8 of [GGZ18]. From (A.4) of [GGZ18], it holds that

[TABLE]

Using the notations in their Lemma 8, we denote

[TABLE]

then the following relations hold

[TABLE]

Taking $j=0$ in (50) gives

[TABLE]

Therefore, by (49) we obtain for $nT\leq t<(n+1)T,n\in\mathbb{N}$

[TABLE]

Then the processes in (34) is uniformly bounded in $L^{2}$ by (48) and (51),

[TABLE]

and

[TABLE]

Similarly, from (50) and (51), we obtain for $nT\leq k<(n+1)T,n\in\mathbb{N}$ ,

[TABLE]

and the upper bounds for $\sup_{k\in\mathbb{N}}E[\|\tilde{V}^{\lambda}_{k}\|^{2}],\sup_{k\in\mathbb{N}}E[\|\tilde{X}^{\lambda}_{k}\|^{2}]$ are $C^{a}_{v},C^{a}_{x}$ , respectively. ∎

Lemma 5.4.

Let $\mu,\nu$ be any two probability measures on $\mathbb{R}^{2d}$ . It holds that

[TABLE]

where $c_{17}:=3\max\{1+\alpha_{c},\gamma^{-1}\}$ .

Proof.

From (2.11) of [EGZ17], we have that $h(x)\leq x$ , for $x\geq 0$ , and from (18), $r((x_{1},v_{1}),(x_{2},v_{2}))\leq c_{17}/3\|(x_{1},v_{1})-(x_{2},v_{2})\|$ . By definition (20), we estimate

[TABLE]

∎

Lemma 5.5.

Let $1\leq q\in\mathbb{N}$ . It holds that

[TABLE]

Proof.

We will use the arguments in the proof of Lemma 12 of [GGZ18] to obtain the contraction for $\mathcal{V}(X^{\lambda}_{k},V^{\lambda}_{k})$ and in Lemma 3.9 of [CMR*+*18] to obtain high moment estimates. First, we have

[TABLE]

Denoting $\Delta^{1}_{k}=V^{\lambda}_{k}-\lambda[\gamma V^{\lambda}_{k}+g(X^{\lambda}_{k},U_{\mathbf{z},k})]$ , we compute

[TABLE]

Similarly, we have

[TABLE]

Denoting $\Delta^{2}_{k}=X^{\lambda}_{k}+\gamma^{-1}V^{\lambda}_{k}-\lambda\gamma^{-1}g(X^{\lambda}_{k},U_{\mathbf{z},k})$ , we compute that

[TABLE]

Let us denote $\mathcal{V}_{k}=\mathcal{V}(X^{\lambda}_{k},V^{\lambda}_{k})$ . From (52), (LABEL:eq:V), (54) and (55) we compute that

[TABLE]

where

[TABLE]

Using the inequality (16), we obtain

[TABLE]

The quantity $\mathcal{E}_{k}$ is bounded as follows

[TABLE]

As in [GGZ18], we deduce that

[TABLE]

And then we get that

[TABLE]

where

[TABLE]

Similarly, we bound $\Sigma_{k}$ , using (5) and the definitions of $\Delta^{1}_{k},\Delta^{2}_{k}$ ,

[TABLE]

and thus

[TABLE]

where

[TABLE]

Noting that $\lambda_{c}\leq 1/4,$ we have

[TABLE]

From (57), (54) we obtain

[TABLE]

Therefore, for $0<\lambda<\frac{\gamma\lambda_{c}}{2K_{1}}$

[TABLE]

where

[TABLE]

Define $E_{k}[\cdot]:=E[\cdot|(X^{\lambda}_{k},V^{\lambda}_{k}),\mathbf{Z}=\mathbf{z}]$ . We then compute as follows,

[TABLE]

where the last inequality is due to Lemma A.3 of [CMR*+*18]. Denoting $c_{19}:=\gamma A_{c}+\beta K_{2}+\gamma d,$ we continue

[TABLE]

Clearly we have

[TABLE]

Define

[TABLE]

On $\left\{\mathcal{V}_{k}\geq\tilde{M}_{1}\right\}$ we have

[TABLE]

And thus

[TABLE]

If we choose

[TABLE]

then on $\{V_{k}\geq\tilde{M}\}$ , the second, the third and the fourth term in the RHS of (LABEL:eq:2p2) are bounded by [math] and then

[TABLE]

On $\{\mathcal{V}_{k}<\tilde{M}\}$ , we have

[TABLE]

where $\tilde{N}=2c_{19}q\tilde{M}^{2q-1}+6q(2q-1)2^{2q-3}\beta dP_{1}\tilde{M}^{2q-1}+q(2q-1)2^{4q-3}\beta^{q}P^{q}_{1}E\|\xi_{k+1}\|^{2q}\tilde{M}^{q}$ . For sufficiently small $\lambda$ , we get from these bounds

[TABLE]

The proof is complete by using (5). ∎

5.1 Explicit dependence of constants on important parameters

Similar to [GGZ18], we choose $\mu_{0}$ in such a way that

[TABLE]

Then we get $C^{c}_{x}=C^{c}_{v}=C^{a}_{x}=C^{a}_{v}=\mathcal{O}((\beta+d)/\beta).$ It follows that

[TABLE]

It is checked that

[TABLE]

and

[TABLE]

The constant $c_{*},C_{*}$ are $\mu_{*}$ and $C$ respectively in [GGZ18]. In addition, we check

[TABLE]

and hence

[TABLE]

From Lemma 16 of [GGZ18], we get

[TABLE]

Furthermore, it is observed that

[TABLE]

Therefore, for a fixed $k$ , the term $\mathcal{B}_{1}$ is bounded by

[TABLE]

Since $c_{*}$ is exponentially small in $(\beta+d)$ , our bound for $\mathcal{B}_{1}$ is worse than that of $\mathcal{J}_{1}(\varepsilon)+\overline{\mathcal{J}}_{0}(\varepsilon)$ given in [GGZ18].

Acknowledgments

Both authors were supported by the NKFIH (National Research, Development and Innovation Office, Hungary) grant KH 126505 and the “Lendület” grant LP 2015-6 of the Hungarian Academy of Sciences. The authors thank Minh-Ngoc Tran for helpful discussions.

Bibliography35

The reference list from the paper itself. Each links out to its DOI / PubMed record.

1[AP 16] Hedy Attouch and Juan Peypouquet. The rate of convergence of Nesterov’s accelerated forward-backward method is actually faster than 1/k^2. SIAM Journal on Optimization , 26(3):1824–1834, 2016.
2[BBLG 17] Michael Betancourt, Simon Byrne, Sam Livingstone, and Mark Girolami. The geometric foundations of Hamiltonian Monte Carlo. Bernoulli , 23(4A):2257–2298, 2017.
3[BCM + 18] M. Barkhagen, N. H. Chau, E. Moulines, M. Rásonyi, S. Sabanis, and Y. Zhang. On stochastic gradient Langevin dynamics with stationary data streams in the logconcave case. preprint, ar Xiv:1812.02709 , 2018.
4[Bet 17] Michael Betancourt. A conceptual introduction to Hamiltonian Monte Carlo. ar Xiv preprint ar Xiv:1701.02434 , 2017.
5[BV 05] François Bolley and Cédric Villani. Weighted Csiszár-Kullback-Pinsker inequalities and applications to transportation inequalities. In Annales de la Faculté des sciences de Toulouse: Mathématiques , volume 14, pages 331–352, 2005.
6[CCAY + 18] Xiang Cheng, Niladri S. Chatterji, Yasin Abbasi-Yadkori, Peter L. Bartlett, and Michael I. Jordan. Sharp convergence rates for Langevin dynamics in the nonconvex setting. ar Xiv preprint ar Xiv:1805.01648 , 2018.
7[CCBJ 17] Xiang Cheng, Niladri S Chatterji, Peter L Bartlett, and Michael I Jordan. Underdamped Langevin MCMC: A non-asymptotic analysis. ar Xiv preprint ar Xiv:1707.03663 , 2017.
8[CCG + 16] Changyou Chen, David Carlson, Zhe Gan, Chunyuan Li, and Lawrence Carin. Bridging the gap between stochastic gradient MCMC and stochastic optimization. In Artificial Intelligence and Statistics , pages 1051–1060, 2016.

TL;DR

Contribution

Findings

Abstract

Peer Reviews

Videos

Taxonomy

Abstract

1 Introduction

2 Asumptions and main results

Assumption 2.1**.**

Assumption 2.2**.**

Assumption 2.3** (Dissipative).**

Assumption 2.4**.**

Assumption 2.5**.**

Assumption 2.6**.**

Remark 2.7**.**

Theorem 2.8**.**

Proof.

Theorem 2.9**.**

Proof.

Corollary 2.10**.**

Proof.

3 Related work and our contributions

4 Proofs

4.1 A contraction result

Theorem 4.1**.**

Proof.

4.2 Proof of Theorem 2.8

Proof.

Remark 4.2**.**

Lemma 4.3**.**

Proof.

Lemma 4.4**.**

Proof.

4.3 Proof of Theorem 2.9

4.3.1 The first term T1\mathcal{T}_{1}T1​

4.3.2 The second term T2\mathcal{T}_{2}T2​

Remark 4.5**.**

4.3.3 The third term T3\mathcal{T}_{3}T3​

5 Technical lemmas

Lemma 5.1**.**

Proof.

Lemma 5.2**.**

Proof.

Lemma 5.3**.**

Proof.

Lemma 5.4**.**

Proof.

Lemma 5.5**.**

Proof.

5.1 Explicit dependence of constants on important parameters

Acknowledgments

Assumption 2.1.

Assumption 2.2.

Assumption 2.3 (Dissipative).

Assumption 2.4.

Assumption 2.5.

Assumption 2.6.

Remark 2.7.

Theorem 2.8.

Theorem 2.9.

Corollary 2.10.

Theorem 4.1.

Remark 4.2.

Lemma 4.3.

Lemma 4.4.

4.3.1 The first term $\mathcal{T}_{1}$

4.3.2 The second term $\mathcal{T}_{2}$

Remark 4.5.

4.3.3 The third term $\mathcal{T}_{3}$

Lemma 5.1.

Lemma 5.2.

Lemma 5.3.

Lemma 5.4.

Lemma 5.5.