Skip to content

Architectures for mathematical data

Aug 2026 · AI in Mathematics and Theoretical Physics · pp. 1-9 · 0 citations

TL;DR

This work develops the principle that choosing a neural architecture is not a matter of choosing the largest model, but of choosing an inductive bias that matches the structure the data already has through three self-contained case studies.

Abstract

Mathematical objects rarely arrive as unstructured feature vectors. They arrive as sequences indexed by primes, as directed graphs closed under a combinatorial move, as functions constrained by a differential equation. Choosing a neural architecture is therefore not a matter of choosing the largest model, but of choosing an inductive bias that matches the structure the data already has. We develop this principle through three self-contained case studies, each implemented from scratch in JAX and each runnable on a laptop CPU. (i) Sequential data: a 1D CNN and a single-block transformer predict the rank of an elliptic curve [Formula: see text] from its normalized Frobenius traces [Formula: see text]; a saliency analysis tracked across training shows the network localizing the discriminative signal in the small primes, where the murmuration phenomenon lives. (ii) Graph data: a directed graph isomorphism network classifies quivers by Dynkin mutation type and is permutation equivariant to machine zero by construction; trained only on quivers with [Formula: see text] and [Formula: see text] vertices, it transfers to quivers on [Formula: see text] vertices. (iii) Continuous data: a physics-informed neural network solves a two-point boundary value problem, after which interval arithmetic bounds the residual rigorously over the entire domain rather than at the collocation points. Based on a tutorial delivered at DANGER: Data, Numbers, and Geometry (Banff, April 2026).

View source

Similar papers

#graph neural networks Preprint Aug 2026

Learning Topological Features of $\widehat Z$-invariants

This paper initiates a systematic approach to handling mathematical data structured as (truncated) infinite $q-series, or equivalently, infinite series of integers, and demonstrates that neural networks can reliably extract essential topological information, such as homology class and underlying graph structure, direct...

Brandon Robinson, Shimal Harichurn, Fabian Ruehle et al. · 0 citations
#machine learning Preprint Sep 2026

Hidden Activations are not Enough I: Knowledge Matrices as Higher Representations

We study the knowledge matrix of a trained feedforward network as a higher representation of its inputs. A network is a pair $(W,f)$, a thin representation $W$ of its quiver and an activation $f$; its function factorizes through the space of quiver representations, each input $x$ inducing a representation, and the know...

Marco Armenta · 0 citations
#graph neural networks Open access Sep 2026

Distance-Based Representation Learning with Nonlinear Feature Algebras

<jats:p> Graph neural networks and spectral embeddings aggregate local neighbourhoods and so miss the global metric properties—growth rate, hyperbolicity, boundary at infinity—that govern large-scale structure in hierarchical, networked, and negatively curved data. We propose a fr...

K. Enakoutsa · 0 citations
#machine learning Preprint Sep 2026

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such a...

Shuang Liang, Le-Jun Liao, Shi-Yuan Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Canonical locks that encode part-whole hierarchies

One of the challenges in representational learning is how to encode part-whole hierarchies in a neural net. Prior works rely on flattening tree-like structures into string-like sequences and training a sequence-to-sequence model via autoregression. While such a representation works for parse-trees in NLP, it is not ent...

Rajat Modi, Y. Rawat · 0 citations
#natural language process... Preprint Aug 2026

All You Need Is Non-Commutative Words

It is shown that the noncommutativity of matrix product captures word order without positional encodings (PEs) and yields several capabilities, including antisymmetric self-attention with no query, key, or value projections, and parallel composition of variable-length text chunks at a reduced attention cost.

Carla Quispe Flores, Stanley Salvatierra, Renan Cabrera · 0 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.