Stable FP4 Training via Transposition-Invariant Block Quantization
This work proposes a low-precision training framework based on 2D block FP4 quantization, which enforces transposition-invariant scaling and preserves consistency between forward and backward computations, and combines this with truncation-free scaling and stochastic rounding to control quantization error and maintain unbiased gradients.