Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets
It is established that wide residual networks (ResNets) with constant scaling factors become asymptotically unlearnable as depth increases, and the generalization capability of wide ResNets can be approximated by kernel regression associated with the Neural Tangent Kernel (NTK).