Skip to content

Sharp Stability Threshold and Certification for Designing Stable Residual Architectures

Jul 2026 · arXiv.org · Vol abs/2607.14576 · 0 citations · 55 references
Computer Science Mathematics

Abstract

We propose \emph{the sublinear-growth principle} for deep residual architectures -- a sharp stability threshold on the input-magnitude exponent of every residual block's velocity field: $$\|v(x, t)\| \leq c\,\|x\|^q + b, \qquad q \in [0, 1].$$ The threshold $q = 1$ is established via two independent arguments. Classical ODE theory gives a global forward flow on $[0, T]$ at $q \le 1$ and exhibits divergent velocity fields at any $q>1$. The optimal-control analysis, via the Hamilton-Jacobi-Bellman equation, sharpens this to a selection statement: the training optimum is bang-bang on the boundary of the admissible class, so the optimum at $q>1$ blows up while the optimum at $q \le 1$ is safe by construction. The exponent criterion $q \le 1$ is thereby a necessary and sufficient condition for stable training. It clarifies architectural placements that ensure the stability of training and inference, explaining, for instance, the stabilizing role of layer normalization. The sublinear-growth velocity fields form \emph{the right function space} on which forward dynamics, adjoint sensitivity, and architectural composition are all well-controlled. An arithmetic of input-magnitude exponents under the five operations that build residual blocks enables efficient certification of $q_k \le 1$ at the level of architectural primitives, in place of ad hoc trial and error in the search for stable neural architectural designs. A parameter-free modification reduces the supercritical Mamba block from $q = 5$ to $q = 1$ without layer normalization, demonstrating this point. Experiments on Mamba and PatchTST confirm that the $q \le 1$ variants train stably: the criterion is the input-magnitude exponent, not the presence of a normalization layer.

View source

Similar papers

#machine learning Preprint Oct 2026

Universal interpolation for deep residual self-attention networks

Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consi...

Sibylle Marcotte, Joan Bruna · 0 citations
Preprint Aug 2026

DeepGOF-1: A Pretrained Convolutional Goodness-of-Fit Test for Logistic Regression with a Computable Consistency Certificate

Goodness-of-fit tests for logistic regression are least reliable where they are most needed: at small samples their levels drift from the nominal one, and combining them worsens the drift. We propose a test whose statistic is a convolutional network, trained once on simulated departures, that reads misfit as a picture:...

Ebrahim Khaled Ebrahim · 0 citations
Preprint Aug 2026

Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth

This work gives a sufficient condition for depth-safety: once an operator's per-step displacement is small relative to the decoder margin, the decoded answer cannot change under further iterations, and gives four operational criteria for useful test-time depth.

Ivan Viakhirev, Kirill Borodin, Amirah Almutairi et al. · 5 citations
Preprint Aug 2026

Training-Free Universal Approximation by Prompting Random Transformers

How expressive is prompting a transformer? Answering this question is important for separating the roles of prompting, architecture, and pretraining in transformer models, and for determining whether task-specific behavior must be stored in model weights or can instead be induced at inference time through the prompt. W...

Alexander Hsu, Rong-Jie Lai · 1 citation
Preprint Aug 2026

TESLA: Taylor Expansion of Sinusoidal Learnable Activations

TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components is proposed, indicating that activation-level degree control transfers to more general vision workloads.

Daehwa Ko, Jae-Hwan Kim, Seunghyun Ham et al. · 0 citations
Open access Sep 2026

DWAT: Density-Weighted Adversarial Training for Robustness Beyond the Training Perturbation Budget

Deep neural networks (DNNs) are widely deployed in safety-critical applications such as medical diagnosis and autonomous driving. Adversarial training (AT) is among the most effective defenses, casting robust optimization as a min–max problem over a defender-specified ℓp-ball of fixed radius ϵ. Bounded defenses of this...

Jie-Ying Huang, Rui-Ming Zhu, Jia Xu et al. · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.