Preprint
Aug 2026
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
VLZip is introduced, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer, and establishes an efficient and powerful new standard for long-context multimodal AI.
Yuqi Zhang, Cheng Chen, Yuyu Guo et al.
· 0 citations