Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks

ML NNGTB POA
我们研究了最速下降算法的一般家族在深度同质神经网络中的隐式偏差,这些算法包括梯度下降、符号下降和坐标下降。我们证明,一旦网络达到完美的训练精度,一种依赖于算法的几何裕度就开始增加,并且表征了算法在后期阶段的偏差。特别地,我们定义了一种优化问题的广义平稳性概念,并展示了这些算法逐步减少了一种(广义的)Bregman散度,这种散度量化了接近最大裕度问题的平稳点的程度。然后,我们通过实验详细观察了使用各种最速下降算法优化的神经网络的轨迹,突显了与Adam算法隐式偏差的联系。
We study the implicit bias of the general family of steepest descent algorithms, which includes gradient descent, sign descent and coordinate descent, in deep homogeneous neural networks. We prove that an algorithm-dependent geometric margin starts increasing once the networks reach perfect training accuracy and characterize the late-stage bias of the algorithms. In particular, we define a generalized notion of stationarity for optimization problems and show that the algorithms progressively reduce a (generalized) Bregman divergence, which quantifies proximity to such stationary points of a margin-maximization problem. We then experimentally zoom into the trajectories of neural networks optimized with various steepest descent algorithms, highlighting connections to the implicit bias of Adam.
许愿