Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction

While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pixel-aligned approaches suffer from spatial inflexibility and massive structural redundancy, whereas query-based methods lack 3D priors and entangle geometry with appearance, yielding blurry, pose-dependent results. To overcome these deficiencies, we propose \textbf{QuerySplat}, a feed-forward 3DGS framework driven by geometric priors and explicit appearance decoupling. Specifically, we design a dual-branch query-based decoder: the geometry branch leverages a pretrained Vision Geometric Model for spatial understanding, which intrinsically endows QuerySplat with pose-free modeling capabilities, while the appearance branch recovers high-frequency details through a dedicated pathway separated from geometric attribute regression. Extensive experiments demonstrate that QuerySplat mitigates the blurry rendering issues of early query-based models and consistently outperforms pixel-aligned approaches in rendering fidelity. On the challenging DL3DV benchmark, it achieves state-of-the-art novel view synthesis performance, with average PSNR gains of 2.30 dB and 1.04 dB over the best pose-free and pose-required baselines, respectively. Project Page: https://inspatio.github.io/querysplat.

Yinglong Li, Donghui Shen, Xiaoyu Zhang et al. · 0 citations
Aug 2026

Flexible Motion Stylization via Multi-modality Latent Diffusion Model.

Multi-modality motion stylization presents a solution to the challenge of generating flexible, stylized motion based on multimodal style inputs. Historically, motion stylization has grappled with the difficulty of balancing content and style, often prioritizing one at the expense of the other. This paper addresses the complex challenge of multi-modality content-style duality, achieving a sophisticated integration that both preserves and enhances the core narrative through nuanced stylistic modifications. We propose a Multi-modality Latent Diffusion Model (MM-LDM), a novel framework that leverages diffusion models under multi-modality conditions, including motion style (text-based or motion-based), motion content, and motion trajectory components. A central innovation in our approach is the introduction of a Multi-condition Denoiser, which carefully balances the preservation of primary content with the dynamic integration of style and trajectory as secondary conditions. This multi-modality guidance mechanism, implemented during the denoising process, ensures that new styles are seamlessly integrated with the original content. It gives rise to more authentic and cohesive motion stylization outcomes, establishing a new benchmark in computer animation. To further refine the control over the text-based motion style, we introduce an LLM parser that converts broad motion descriptions into detailed, part-specific representations. By decomposing the human body into movement-related parts, our method significantly enhances the precision and effectiveness of text-based motion stylization, enabling fine-grained control over individual body parts. Our model's effectiveness and generalization capabilities have been rigorously validated through extensive experiments, including text-based motion stylization and generating stylized motion with video sources, which all demonstrate the potential of our MM-LDM to advance the state-of-the-art motion stylization.

Wenfeng Song, Xingliang Jin, Shuai Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.