News
Layer-Wise Generalized Smoothness: A Unified Framework for Adaptive LMO Optimizers
Abstract
Recent algorithms based on the Linear Minimization Oracle (LMO) framework, such as Muon and Scion, are emerging as viable replacements for the Adam optimizer. They offer improved memory efficiency, better hyperparameter transferability, and superior empirical performance in large-scale tasks like large language model (LLM) training. However, a significant gap remains between their practical success and theoretical understanding. To address these issues, we propose a generalized layer-wise LMO framework alongside a refined generalized smoothness model. This approach accurately captures the layer-wise geometry of neural networks, yielding convergence guarantees with strong predictive power. Experiments with NanoGPT and CNN confirm that our assumption holds along the optimization trajectory, ultimately closing the gap between theory and practice.
Keywords
Edition
Proceedings of the Institute for System Programming, vol. 38, issue 4, part 1, 2026, pp. 25-48
ISSN 2220-6426 (Online), ISSN 2079-8156 (Print).
DOI: 10.15514/ISPRAS-2026-38(4)-2
For citation
Full text of the paper in pdf (in Russian)
Back to the contents of the volume