Skip to content

Verified Deep Learning with Lean 4

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

A formal verification of the math of modern deep learning, written as a Lean 4 blueprint project and shipped as the code it proves. Mathlib's fderiv is the foundation; every layer's backward is derived as a vector-Jacobian product rather than asserted -- dense, ReLU, softmax cross-entropy, 2D convolution, max-pool, BatchNorm's three-term backward, residual and bottleneck blocks, depthwise and inverted residuals, squeeze-excitation, the Universal Inverted Bottleneck, LayerNorm, and a 30-theorem treatment of multi-head self-attention -- and composed up to whole-network VJPs. Nine chapters climb from a linear classifier on MNIST to the Vision Transformer, and a Bestiary chapter maps the same primitives across vision classifiers, semantic segmentation, image generation, reinforcement learning, NLP, diffusion, and beyond-vision domains (audio, 3D scene reconstruction, scientific). Each chapter's network is a Lean 4 NetSpec whose train step -- forward, loss, backward and optimizer in one StableHLO graph -- is rendered from the proofs, with every parameter update tied to the certified gradient at the denotational level, for all seven networks (ResNet-34, ResNet-50, MobileNetV2, MobileNetV4-Conv-M, EfficientNet-B0, ConvNeXt-T, ViT-Tiny) at the artifact that trains. The rendered graphs train end to end through XLA/PJRT, with IREE as a second, independent lowerer for the differential oracle against a parallel JAX pipeline, in four tiers: MNIST, CIFAR-10, Imagenette and full ImageNet-1k, where the verified path lands within 0.3 of its JAX reference on every net run (ResNet-50 at 77.98% top-1 on the RSB-A3 recipe, ConvNeXt-T at 81.30%). Demos extend the same stack, in the Bestiary chapter's order, to drone-view and industrial object detection, sign-language and leaf-disease classification under honest splits, brain-tumour segmentation, a DQN scored against blackjack's exact policy, character-level language models, diffusion, a flow-matching Boltzmann generator, gravitational-wave detection on LIGO strain against the matched filter, and a transformer wavefunction for the Ising chain. The proof suite contains zero project axioms; every theorem closes under exactly {propext, Quot.sound, Classical.choice}, with 73 headline theorems independently kernel-rechecked by leanprover/comparator. Builds on the architectural coverage of 'Convolutional Neural Networks with Swift for TensorFlow' (Apress, 2021), extending through ConvNeXt and ViT and adding formal proofs throughout.

View source

Similar papers

#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11
#computer vision Open access Sep 2012

Making the leap to a software platform strategy: Issues and challenges

A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.

Yaser Ghanam, F. Maurer, P. Abrahamsson · 41 citations · ⚡3
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.

Yang Yue, Shu Li, Yihua Cheng et al. · 15 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.

Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al. · 13 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.