Certifying Residual Architectures from Their Primitives: A Sharp Stability Threshold
TL;DR
The results provide a way to assess the trade-off between stability and representational flexibility directly from block primitives, before training, and improve out-of-distribution generalization in operator learning and accuracy in time-series forecasting.
Abstract
Whether a deep residual architecture trains stably is usually determined by training it, which is expensive and answers the question only for the architecture, depth, and floating-point format tested. In practice, stability is secured by heuristics for where to place normalization and which type to use, supported by experiments but lacking a common principle. We show that stability can instead be certified before training, directly from the architectural primitives of a residual block, as an explicit function of depth and floating-point format. The certificate has two elements: (a) a power-law growth bound, $|v(x)|\le c|x|^q+b$, whose exponent $q$ is computed from the block's primitives by an arithmetic of exponents (forward); and (b) gradient bounds from a Lipschitz condition on the block over reachable states (backward). The certificate determines when the state can reach the largest finite floating-point value $M$: for $q\le1$, overflow requires at least order $\log M$ layers, whereas for $q>1$ it can occur within order $\log_q\log M$ layers; the threshold $q=1$ is sharp. Every normalization and bounded activation sets $q=0$, and the arithmetic identifies minimal modifications that bring a block from $q>1$ to $q=1$ without normalization. Empirically, in forecasting, only $q>1$ blocks blow up, as predicted by the forward certificate. In GPT-2 on OpenWebText, $q=1$ blocks without normalization also diverge: the forward certificate holds, while the backward one fails in the attention block with the largest query-key coefficient. Relaxing normalization from $q=0$ to $q=1$ improves out-of-distribution generalization in operator learning and accuracy in time-series forecasting. As deep models become more complex and costly to train, our results provide a way to assess the trade-off between stability and representational flexibility directly from block primitives, before training.