Skip to content
Preprint

Certifying Residual Architectures from Their Primitives: A Sharp Stability Threshold

Jul 2026 · 0 citations · 32 references
Computer Science Mathematics

TL;DR

The results provide a way to assess the trade-off between stability and representational flexibility directly from block primitives, before training, and improve out-of-distribution generalization in operator learning and accuracy in time-series forecasting.

Abstract

Whether a deep residual architecture trains stably is usually determined by training it, which is expensive and answers the question only for the architecture, depth, and floating-point format tested. In practice, stability is secured by heuristics for where to place normalization and which type to use, supported by experiments but lacking a common principle. We show that stability can instead be certified before training, directly from the architectural primitives of a residual block, as an explicit function of depth and floating-point format. The certificate has two elements: (a) a power-law growth bound, $|v(x)|\le c|x|^q+b$, whose exponent $q$ is computed from the block's primitives by an arithmetic of exponents (forward); and (b) gradient bounds from a Lipschitz condition on the block over reachable states (backward). The certificate determines when the state can reach the largest finite floating-point value $M$: for $q\le1$, overflow requires at least order $\log M$ layers, whereas for $q>1$ it can occur within order $\log_q\log M$ layers; the threshold $q=1$ is sharp. Every normalization and bounded activation sets $q=0$, and the arithmetic identifies minimal modifications that bring a block from $q>1$ to $q=1$ without normalization. Empirically, in forecasting, only $q>1$ blocks blow up, as predicted by the forward certificate. In GPT-2 on OpenWebText, $q=1$ blocks without normalization also diverge: the forward certificate holds, while the backward one fails in the attention block with the largest query-key coefficient. Relaxing normalization from $q=0$ to $q=1$ improves out-of-distribution generalization in operator learning and accuracy in time-series forecasting. As deep models become more complex and costly to train, our results provide a way to assess the trade-off between stability and representational flexibility directly from block primitives, before training.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.