Model Evolution Under Zeroth-Order Optimization: A Neural Tangent Kernel Perspective

Chen Zhang; Yuxin Cheng; Chenchen Ding; Shuqi Wang; Jingreng Lei; Runsheng Yu; Yik-Chung WU; Ngai Wong

arXiv:2603.21169·cs.LG·March 24, 2026

Model Evolution Under Zeroth-Order Optimization: A Neural Tangent Kernel Perspective

Chen Zhang, Yuxin Cheng, Chenchen Ding, Shuqi Wang, Jingreng Lei, Runsheng Yu, Yik-Chung WU, Ngai Wong

PDF

Open Access

TL;DR

This paper introduces the Neural Zeroth-order Kernel (NZK) to analyze the training dynamics of zeroth-order optimization in neural networks, providing theoretical insights and empirical validation for model evolution and convergence acceleration.

Contribution

It develops the NZK framework to characterize ZO training dynamics, extending NTK theory to zeroth-order methods and demonstrating potential for faster convergence.

Findings

01

Expected NZK remains constant during training for linear models

02

Explicit formula for model evolution under squared loss

03

Empirical results show acceleration with shared random vectors

Abstract

Zeroth-order (ZO) optimization enables memory-efficient training of neural networks by estimating gradients via forward passes only, eliminating the need for backpropagation. However, the stochastic nature of gradient estimation significantly obscures the training dynamics, in contrast to the well-characterized behavior of first-order methods under Neural Tangent Kernel (NTK) theory. To address this, we introduce the Neural Zeroth-order Kernel (NZK) to describe model evolution in function space under ZO updates. For linear models, we prove that the expected NZK remains constant throughout training and depends explicitly on the first and second moments of the random perturbation directions. This invariance yields a closed-form expression for model evolution under squared loss. We further extend the analysis to linearized neural networks. Interpreting ZO updates as kernel gradient descent…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStochastic Gradient Optimization Techniques · Neural Networks and Reservoir Computing · Machine Learning and ELM