#natural language process...
Jul 2026
Studying quantization trade-offs for efficient inference deployment in machine translation
The experiments show that the trade-off between inference efficiency and translation quality depends not only on the quantization format, but also on the choice of text chunking strategy, as well as on the choice of text chunking strategy.
Jim Zhao, Sohir Maskey, Koen Oostermeijer et al.
· arXiv.org · 0 citations