Layer-Wise Generalized Smoothness: A Unified Framework for Adaptive LMO Optimizers


Layer-Wise Generalized Smoothness: A Unified Framework for Adaptive LMO Optimizers

Riabinin A.A. (BRAIn Lab, Moscow, Russia)

Abstract

Recent algorithms based on the Linear Minimization Oracle (LMO) framework, such as Muon and Scion, are emerging as viable replacements for the Adam optimizer. They offer improved memory efficiency, better hyperparameter transferability, and superior empirical performance in large-scale tasks like large language model (LLM) training. However, a significant gap remains between their practical success and theoretical understanding. To address these issues, we propose a generalized layer-wise LMO framework alongside a refined generalized smoothness model. This approach accurately captures the layer-wise geometry of neural networks, yielding convergence guarantees with strong predictive power. Experiments with NanoGPT and CNN confirm that our assumption holds along the optimization trajectory, ultimately closing the gap between theory and practice.

Keywords

adaptive optimization; linear minimization oracle; layer-wise smoothness; large language models.

Edition

Proceedings of the Institute for System Programming, vol. 38, issue 4, part 1, 2026, pp. 25-48

ISSN 2220-6426 (Online), ISSN 2079-8156 (Print).

DOI: 10.15514/ISPRAS-2026-38(4)-2

For citation

Riabinin A.A. Layer-Wise Generalized Smoothness: A Unified Framework for Adaptive LMO Optimizers. Proceedings of the Institute for System Programming, vol. 38, issue 4, part 1, 2026, pp. 25-48 DOI: 10.15514/ISPRAS-2026-38(4)-2.

Full text of the paper in pdf (in Russian) Back to the contents of the volume