Skip to content
Preprint

Improved Convergence Rates for Stochastic Multi-Gradient Descent which Close the Gap: A Proof by AI

Jul 2026 · 0 citations · 30 references
Mathematics

TL;DR

A new convergence rate for SMG in terms of the squared Pareto-stationarity (PS) measure is established, to exploit the Lipschitz continuity of the PS measure, defined by the norm of the multi-gradient descent algorithm (MGDA) direction, rather than the $(1/2)-H\"older continuity of the MGDA direction.

Abstract

For smooth nonconvex stochastic multi-objective optimization, stochastic multi-gradient descent (SMG) computes an approximate steepest common descent direction of the objectives from stochastic gradients. Under standard assumptions, this note establishes a new convergence rate for SMG in terms of the squared Pareto-stationarity (PS) measure. With a constant stepsize and linearly growing mini-batches, the expected squared empirical PS measure at the algorithm's output is $\tilde{O}(T^{-1})$ after $T$ iterations. This improves on the $\tilde{O}(T^{-1/4})$ bound obtained by Chen et al. (JMLR, 2024) for linearly growing batches, without requiring bounded gradients. Here, $\tilde{O}(\cdot)$ suppresses logarithmic factors. The key to the rate improvement is to exploit the Lipschitz continuity of the PS measure, defined by the norm of the multi-gradient descent algorithm (MGDA) direction, rather than the $(1/2)$-H\"older continuity of the MGDA direction used by Chen et al. (2024). For variants of MoCo and MoCo+ with exact MGDA updates and constant batch sizes, the analysis gives expected squared empirical PS rates of $O(T^{-1/2})$ and $\tilde{O}(T^{-2/3})$, respectively. These bounds match the corresponding single-objective momentum rates up to logarithmic factors. The proof was accidentally discovered while the author was preparing homework for a graduate course: ChatGPT 5.4 Thinking Extended generated the initial proof strategy for the main SMG result in response to an author-written homework-solution prompt; the author then verified and reorganized the resulting argument. The appendices document the prompt, summarize the student submissions with LLM disclosures, and present additional proofs and results on momentum-based stochastic MGDA.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.