Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, substantially improving long-horizon reliability, and makes three contributions to an environment-grounded training stack with in-context demonstration learning.
Zihan Ding, Longxu Dou, Qixiao Gao et al.· 0 citations
This thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations, and addresses long-horizon modeling through architectures with memory.
LongDS is introduced, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget.