TextPose: Language-Conditioned Static 3D Human Pose Synthesis via Discrete Pose Tokens
Abstract
This paper addresses static text-to-pose synthesis. Given a natural language description of a body configuration, the task is to generate the corresponding single-frame 3D human pose as a set of J root-relative joint positions in ℝJ×3. Discrete latent representations have been applied successfully to text-to-motion generation, where a sequence of frames is compressed into a short sequence of tokens. Whether the same principle helps in the static, single-frame regime is an open question, because a single pose offers no temporal redundancy to compress. We study this question with TextPose, a two-stage framework. In the first stage, a Vector Quantized Variational Autoencoder (VQ-VAE) with a bone-length regularisation loss encodes each pose as a sequence of N discrete tokens drawn from a learned codebook of size K; we use N = 8 and K = 512, giving a genuine compression of the 66-dimensional pose vector. In the second stage, a frozen RoBERTa encoder represents the description, and a Transformer decoder conditioned by cross-attention at every layer autoregressively predicts the token sequence; the first-stage decoder then recovers the 3D joints. We construct a static-pose benchmark from HumanML3D by extracting minimum-velocity keyframes, filtering action-only descriptions, and splitting at the level of source clips to preclude pose leakage. We compare against a published static baseline, a sentence-embedding regressor, and, critically, a parameter-matched continuous variant of our own decoder, which isolates the contribution of discretisation from that of model capacity. We pre-specify trivial floors, oracle ceilings, diversity measures, and a paired-bootstrap test so that the eventual comparison is interpretable and cannot be chosen after the results are observed.