Skip to content

Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

This work introduces Schur Replay, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns.

Abstract

Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosing that scale is difficult in GPTQ because quantizing one column updates those that follow, so evaluating a block independently can misestimate its final reconstruction error. Large models pose a second challenge: full-precision weights, calibration activations, and second-order state cannot all remain on one accelerator, while assigning complete layers to devices leaves each time-consuming layer solve serial. We introduce \emph{Schur Replay}, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns. Separately, our execution infrastructure keeps only the active layer resident, tiers activations across device, host, and disk, retires full-precision layers after export, and distributes independent output rows across tensor-parallel ranks. Together, the algorithm and infrastructure attain $99.35\%$ and $100.84\%$ question-weighted recovery from BF16 across seven benchmarks on Qwen3.5-397B-A17B and Llama-3.3-70B-Instruct. On the 397B model, the infrastructure reduces measured per-layer time by $15.17\times$ over ModelOpt and $23.14\times$ over LLM Compressor, with lower memory used per GPU.

View source

Similar papers

#natural language process... Preprint Aug 2026

H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference

H-Scale is a lightweight post-processing method for NVFP4 per-group scale refinement that selects hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, thereby targeting layer output perturbation more directly.

Hao Yu, Zheng Li, Dayiheng Liu et al. · 1 citation · ⚡1
#artificial intelligence Preprint Sep 2026

Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer

This work builds a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning.

Hao-Ran Jin, Kang-Qi Zhang, Ji-Rong Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

ThinQuant is introduced, a data selection procedure which reduces the required number of calibration data points and an exact reduction of the associated optimization on this reduced calibration set using an efficient ADMM algorithm that iteratively employs thin matrix updates at every step, hence the name ThinQuant.

Mehdi Makni, Ryan Lucas, R. Mazumder · 0 citations
#small language model Preprint Sep 2026

SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models

SPHQuant is a rotation-free spherical weight-only quantization framework for VLMs that isolates outlier magnitude into the radius while keeping directions bounded and statistically regular and designs a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radia...

Ke-Wei Zhang, Zheng Chen, Hao-Tong Qin et al. · 0 citations
Preprint Aug 2026

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

SCHUROPT is introduced, which analytically eliminates the suffix's optimal continuous response, yielding an exact groupwise quadratic with Schur-complement curvature, and achieves the highest mean zero-shot accuracy among the evaluated backpropagation free PTQ baselines.

Gunjun Lee, Sehwan Son, Younjoo Lee et al. · 0 citations
#machine learning Preprint Sep 2026

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

This paper profiles vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL across cold prefill, warm prefill, and decode, sweeping sequence length and batch size, and replaces the usual single utilization number with eight counter-validated views derived from raw Nsight Compute reports.

Mohammad Siavashi, Gerald Q. Maguire, Dejan Kostic et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.