Skip to content
Conference

TextPose: Language-Conditioned Static 3D Human Pose Synthesis via Discrete Pose Tokens

Aug 2026 · 2026 6th International Conference on Computer Science and Blockchain (CCSB) · pp. 670-677 · 0 citations · 16 references

Abstract

This paper addresses static text-to-pose synthesis. Given a natural language description of a body configuration, the task is to generate the corresponding single-frame 3D human pose as a set of J root-relative joint positions in ℝJ×3. Discrete latent representations have been applied successfully to text-to-motion generation, where a sequence of frames is compressed into a short sequence of tokens. Whether the same principle helps in the static, single-frame regime is an open question, because a single pose offers no temporal redundancy to compress. We study this question with TextPose, a two-stage framework. In the first stage, a Vector Quantized Variational Autoencoder (VQ-VAE) with a bone-length regularisation loss encodes each pose as a sequence of N discrete tokens drawn from a learned codebook of size K; we use N = 8 and K = 512, giving a genuine compression of the 66-dimensional pose vector. In the second stage, a frozen RoBERTa encoder represents the description, and a Transformer decoder conditioned by cross-attention at every layer autoregressively predicts the token sequence; the first-stage decoder then recovers the 3D joints. We construct a static-pose benchmark from HumanML3D by extracting minimum-velocity keyframes, filtering action-only descriptions, and splitting at the level of source clips to preclude pose leakage. We compare against a published static baseline, a sentence-embedding regressor, and, critically, a parameter-matched continuous variant of our own decoder, which isolates the contribution of discretisation from that of model capacity. We pre-specify trivial floors, oracle ceilings, diversity measures, and a paired-bootstrap test so that the eventual comparison is interpretable and cannot be chosen after the results are observed.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.