Path-Averaged Gradient Approximation Algorithms for Deep Neural Network Training
Ieee AccessPeer ReviewedRafal Wolniak +22026Magazines
We explore previously unreported properties and practical utility of integrating gradients along parameter trajectories for training deep neural networks. Unlike standard methods that rely on instantaneous gradients, this approach averages gradients over a continuous range of parameter values at each update step. Our contributions are: (a) We show that, across multiple convolutional architectures, this method yields up to 53.5% greater reduction in per-batch loss compared to baseline optimizers. (b) We demonstrate that for a fixed batch, a single step can approximate more than four standard gradient steps in models lacking skip connections and thus prone to ill-conditioned curvature. (c) We introduce an efficient approximation of the average gradient for ResNet-152 fine-tuning that aggregates gradients over hundreds of past training iterations on a fixed batch at each update, yielding improved accuracy across optimizers. We validate these contributions using representative first-order (RMSProp, Adam) and second-order (SOAP) optimizers. These results highlight parameter-space integrated gradients as a novel framework for understanding and improving generalization and local-descent efficiency in deep neural networks.
The content you want is available to Zendy users.
Already have an account? Sign inHaving issues? Contact support