Papers
Topics
Authors
Recent
Search
2000 character limit reached

Accelerated Zeroth-Order and First-Order Momentum Methods from Mini to Minimax Optimization

Published 18 Aug 2020 in math.OC, cs.CV, and cs.LG | (2008.08170v7)

Abstract: In the paper, we propose a class of accelerated zeroth-order and first-order momentum methods for both nonconvex mini-optimization and minimax-optimization. Specifically, we propose a new accelerated zeroth-order momentum (Acc-ZOM) method for black-box mini-optimization where only function values can be obtained. Moreover, we prove that our Acc-ZOM method achieves a lower query complexity of O~(d<sup>3/4ϵ<sup>−3)\tilde{O}(d<sup>{3/4}\epsilon<sup>{-3}) for finding an ϵ\epsilon-stationary point, which improves the best known result by a factor of O(d<sup>1/4)O(d<sup>{1/4}) where dd denotes the variable dimension. In particular, our Acc-ZOM does not need large batches required in the existing zeroth-order stochastic algorithms. Meanwhile, we propose an accelerated zeroth-order momentum descent ascent (Acc-ZOMDA) method for black-box minimax optimization, where only function values can be obtained. Our Acc-ZOMDA obtains a low query complexity of O~((d1+d2)<sup>3/4κy<sup>4.5ϵ<sup>−3)\tilde{O}((d_1+d_2)<sup>{3/4}\kappa_y<sup>{4.5}\epsilon<sup>{-3}) without requiring large batches for finding an ϵ\epsilon-stationary point, where d1d_1 and d2d_2 denote variable dimensions and κy\kappa_y is condition number. Moreover, we propose an accelerated first-order momentum descent ascent (Acc-MDA) method for minimax optimization, whose explicit gradients are accessible. Our Acc-MDA achieves a low gradient complexity of O~(κy<sup>4.5ϵ<sup>−3)\tilde{O}(\kappa_y<sup>{4.5}\epsilon<sup>{-3}) without requiring large batches for finding an ϵ\epsilon-stationary point. In particular, our Acc-MDA can obtain a lower gradient complexity of O~(κy<sup>2.5ϵ<sup>−3)\tilde{O}(\kappa_y<sup>{2.5}\epsilon<sup>{-3}) with a batch size O(κy<sup>4)O(\kappa_y<sup>4), which improves the best known result by a factor of O(κy<sup>1/2)O(\kappa_y<sup>{1/2}). Extensive experimental results on black-box adversarial attack to deep neural networks and poisoning attack to logistic regression demonstrate efficiency of our algorithms.

Citations (49)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.