This work introduces Schur Replay, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns.
Rui-Ying Ding, Jie Li, Kang He et al.· 0 citations
This work formalizes per-batch dispatch as a fixed-charge makespan problem---NP-hard on two fully replicated GPUs, polynomial in degenerate limits---and presents TEMPO, a makespan-aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out-of-process and fuses dispatch with count...
Jie Li, Chen-Xin Jia, Jinliang Shen et al.· 1 citation
A lightweight Burst-Gap model parameterized by backpressure-free issue time, effective drain rate, and outstanding capacity predicts issue overhead, recovery between bursts, and the onset of backpressure, and two communication-computation fused kernels are redesigned.
Jianwen Xian, Zhiyuan Xu, Yucheng Li et al.· arXiv.org· 0 citations
The new version of AlayaWorld substantially revise how conditioning signals are represented and integrated into the model, replacing the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer.
AlayaWorld Team Kaipeng Zhang, Chuan-Hao Li, Yi-Fan Zhan et al.· 5 citations· ⚡2
AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning, and the framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-wit...
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li et al.· arXiv.org· 1 citation
Sekai2 is introduced, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling, and Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant tem...
Kang He, Wen-Shuo Peng, Zi-Hui Gao et al.· 2 citations· ⚡1
AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 s...
AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al.· arXiv.org· 4 citations
HeadCast is proposed, a training-free, plug-and-play acceleration framework built on the observation that a pre-trained AR model's attention heads exhibit stable, heterogeneous behaviors that accelerates inference by up to 1.62x at 720P and 1.95x at 1080P, while keeping VBench quality on par with full attention and lar...
Jinliang Shen, Li Su, Zheming Li et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.