Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation
A unified spatiotemporally decoupled framework named DeMoDiff is proposed, which jointly redesigns representation and architecture and incorporates spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability.