Skip to content

Author

Haotong Qin

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

S $^{2}$ Q-VDiT$^+$: Accurate Quantized Video Diffusion Transformer with Multi-Resolution Sampling and Structural Distillation.

Large-scale video diffusion models (V-DMs) have achieved remarkable text-to-video generation quality, yet their massive computational complexity makes deployment costly. Post-Training Quantization (PTQ) offers an appealing route to accelerate inference without retraining, but existing diffusion PTQ methods remain fragile on modern V-DMs. A key reason is that contemporary V-DMs are intrinsically multi-resolution due to multi-stage training, while most prior PTQ pipelines calibrate at a fixed resolution, causing suboptimal calibration signals and biased distributions under resolution changes. To address this gap, we propose S $^{2}$ Q-VDiT $^+$, a multi-resolution co-design PTQ framework from data, supervision, and quantizer perspectives. First, Denoising-Prior Based Multi-Resolution Sampling constructs resolution-consistent noisy latents by mapping to the clean space and re-noising, together with a trajectory-aware resolution policy across timesteps. Second, Structure-Aware Multi-Resolution Distillation enhances structural alignment via window-wise distillation and transfers resolution-aware spatial dependencies via multi-scale attention distillation. Third, Debiased Modulated Quantization mitigates skewed distributions using asymmetric weight quantization and a fuseable activation debiasing scheme. Extensive experiments on multiple state-of-the-art video generation models demonstrate that S$^{2}$ Q-VDiT$^+$ consistently outperforms strong PTQ baselines under W4A6 and W4A4, delivers up to $2.08\times$ end-to-end speedup, and reduces model storage and inference memory by up to $3.8\times$ and $2.1\times$, respectively.

Weilun Feng, Chuanguang Yang, Haotong Qin et al. · 2 citations
Preprint Aug 2026

Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding

Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and contactless sensing modality, but radar observations are difficult to align with language. Existing radar-language methods often rely on synthetic data or lack explicit supervision for human body structure and motion. We present mmMind, a radar-language model that uses synchronized 3D pose as training-only supervision. A spatio-temporal radar encoder is pretrained to capture body configuration and motion dynamics, after which the pose head is removed so that inference requires radar alone. The learned radar representations are then aligned with an LLM for behavior captioning and spatio-temporal question answering. We also introduce mmMind-Bench, a real-world mmWave-language benchmark containing 17.9 hours of recordings from 23 participants across seven indoor environments. Experiments on captioning, question answering, and unseen-action generalization show that mmMind consistently outperforms existing radar-language baselines, while ablations confirm the importance of pose-guided pretraining.

Duo Zhang, Zhehui Yin, Zhiyun Yao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.