Skip to content

Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

Jul 2026 · arXiv.org · Vol abs/2607.26001 · 0 citations · 36 references
Computer Science Mathematics

TL;DR

Across ImageNet-1K experiments on ViT-Small/16 and ResNet-50, it is found that the combination of a spectral inner step with a Muon outer step performs consistently strongly, achieving the best validation accuracy on both models among the evaluated methods.

Abstract

Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a"small"perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear perspective on which geometries are most effective in practice remains elusive. Recent work on matrix-aware optimization, particularly the Muon optimizer, suggests that respecting the matrix structure of hidden-layer weights can lead to strong empirical performance. Motivated by this, we study matrix-aware geometry in both stages of SAM: we introduce a layerwise spectral inner perturbation for matrix-valued hidden-layer parameters and combine it with either AdamW/SGDW or Muon in the outer update. Across ImageNet-1K experiments on ViT-Small/16 and ResNet-50, we find that the combination of a spectral inner step with a Muon outer step performs consistently strongly, achieving the best validation accuracy on both models among the evaluated methods.

View source

Similar papers

Jul 2026

Gradient-Energy Guided Block-Wise Perturbations for Sharpness-Aware Minimization

Sharpness-Aware Minimization (SAM) improves generalization by minimizing the worst-case loss in a local parameter neighborhood. Standard SAM implicitly allocates its global perturbation budget across parameter blocks according to instantaneous minibatch gradient norms. Such an allocation can be noisy and may not reflec...

Zhenghao Huang, Jiaxin Deng, Junbiao Pang · 0 citations
Conference Open access Aug 2026

EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations

Recent progress in optimization research has highlighted the sharpness of the loss landscape as a key factor in narrowing the generalization gap. Motivated by this insight, Sharpness-Aware Minimization (SAM) was proposed as a training strategy that enhances generalization. Despite the promising performance, SAM suffers...

Tanapat Ratchatorn, Masayuki Tanaka · 0 citations
#machine learning Preprint Sep 2026

Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization

Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences,...

Nour Jamoussi, Marios Kountouris · 0 citations
#small language model Preprint Aug 2026

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

Spectral-Aware Muon is introduced, which holds the head at the Muon scale and amplifies the bulk using a static spectral prior, and both variants outperform tuned AdamW and Muon (Scion implementation) baselines in all evaluated model-scale and batch-size configurations.

Xiaodong Wu, Wenyi Yu, Chao Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.