Skip to content

Author

T. Hoefler

We have 5 of 611 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Gauss What You Need: Compact Gaussian Splatting Across Scene Scales

3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to select this number automatically across capture scales remains un...

Afif Boudaoud, Jia-Yi Liu, A. Calotoiu et al. · 0 citations
Preprint Aug 2026

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate high...

Sherry Xu, M. Heddes, Jackson Peng et al. · 0 citations
Preprint Sep 2026

Quantum computers will not be that different: A blueprint for quantum computer architecture at scale

Quantum computers are technologically novel and unusual, but at system scale they should be engineered using many of the same principles that govern classical heterogeneous accelerators. This paper argues that utility-scale quantum architecture is primarily a cost-performance problem across a coupled quantum-classical...

T. Hoefler, M. Troyer · 0 citations
Preprint Aug 2026

Performance Foundations of Parallel&Distributed Reasoning Language Models

This work systematize the RL-for-LLM paradigm and provides a compute-centric analysis of prominent post-training algorithmic frameworks: Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), as well as their variants, and develops a taxonomy of intra- and inter-model parallelism strategies for...

Maciej Besta, L. Schmidt, Lara Nonino et al. · 0 citations
Jul 2026

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

Building on NCCL's device-side API, low-latency interfaces for constructing custom collective kernels are developed and used to implement new symmetric collectives in NCCL, demonstrating benefits for both AI inference and traditional HPC workloads.

Siyuan Shen, Anton Korzh, J. Bachan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.