Skip to content
Preprint

Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators

Jul 2026 · 0 citations · 36 references
Computer Science

TL;DR

This work presents the first production-scale application of Triton on a custom ML accelerator, MTIA-2i, developed by Meta, and develops a new compiler backend that targets it, introduces enhancements to TorchInductor code generation, and proposes minimal language extensions that expose MTIA-specific architectural features.

Abstract

The rapid growth in machine learning workloads has fueled the proliferation of custom accelerator architectures. Designed from the ground up, these accelerators often expose programming models that are distinct from GPUs. While hyperscalers and AI chip startups continue to innovate in this space, achieving broad operator coverage to support diverse models remains a major challenge. Additionally, an easy-to-use, high-level kernel programming language is important for rapid iteration of models and kernels. Triton, together with TorchInductor, addresses these issues on GPUs, but its viability on accelerators with different programming models has yet to be established. In this work, we present the first production-scale application of Triton on a custom ML accelerator, MTIA-2i, developed by Meta. To support MTIA-2i, we develop a new compiler backend that targets it, introduce enhancements to TorchInductor code generation, and propose minimal language extensions that expose MTIA-specific architectural features. We demonstrate that Triton-MTIA kernels achieve performance competitive with expert-tuned C++ implementations. Leveraging these development efficiency gains, we successfully deployed manually written and Inductor-generated Triton kernels in production across approximately 60 different model types, accounting for 50% of layers and 47% of non-GEMM execution time for these models. Our results provide compelling evidence that DSLs like Triton can bridge the programming model gaps between ML frameworks, kernels, and custom accelerators, enabling rapid innovation and efficient deployment at scale.

View source

Similar papers

Review Aug 2026

AI Hardware Accelerators for Large Language Models: Architectures and the Memory Wall

Large language models (LLMs) place unprecedented and still-growing demands on the hardware that trains and serves them. This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs such as TPUs, Trainium, Groq, and Cerebras, reconfigurable FPGAs, processing-i...

Siddharth Patel, Rohit Singh · 0 citations
Preprint Sep 2026

AccelForge: Comprehensive Modeling and Co-Design Framework for AI Accelerators

Tensor algebra workloads, of which deep neural networks are prominent examples, are energy-intensive workloads in modern datacenter and edge deployments, making accelerators necessary to achieve energy efficiency and high throughput. To quickly evaluate and iterate on accelerator designs, we need an accelerator modelin...

Tanner Andrulis, Michael Gilbert, Vivienne Sze et al. · 0 citations
Book Open access Sep 2026

Scaling the Extended OpenDwarfs: Evaluating Cross-Vendor CUDA Performance Portability with SCALE

CUDA is the dominant GPU programming model in HPC and industrial accelerator software, and a large body of production code is written directly in it. Deploying that code on non-NVIDIA accelerators has traditionally required source translation, backend-specific rewrites, or a full rewrite in a new programming model. Thi...

Beau Johnston, Chris Kitching, Matthew Ireland et al. · 0 citations
Preprint Aug 2026

FSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM Applications

With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. The question of how to efficiently generate optimal AI chip accelerator designs remains unresolved and challenging. Currently, there is a lack of end-to-end design methodol...

J. Mok, Qi-Jun Zhang, Zhi-Yao Xie · 0 citations

Systematic Design Methodologies for Multi-Engine Deep Learning Accelerators

This thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives, and identifies architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy,...

Fareed Mohammad Qararyah · 0 citations
Preprint Sep 2026

FlashGPU-sim: Enabling GPU Modeling for Modern Architectures and AI Workloads

FlashGPU-sim is presented, an open-source, execution-driven, cycle-accurate GPU simulator for modern AI workloads that faithfully models modern hardware features such as asynchronous data movement, fine-grained synchronization, tensor-core execution, and distributed shared memory.

Si-Ying Yu, Yi-Xun Hong, Guo-Zhi Qiu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.