Skip to content

Author

Disheng Liu

We have 6 of 11 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Filling the Unseen: Scene Extrapolation via 3D Gaussian Splatting

3D Gaussian Splatting achieves photorealistic reconstruction within training view distribution, yet it degrades on out-of-distribution novel views, exhibiting holes in unobserved regions and artifacts in observable areas. Recent works formulate this task as extrapolation and interpolation and try to address it with gen...

Yunlai Zhou, Yiren Lu, Tuo Liang et al. · 0 citations
#machine learning Preprint Sep 2026

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visua...

Yan-Yan Zhang, Di-Sheng Liu, Xin-Peng Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models

Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orientation or symmetry-in...

Tuo Liang, Di-Sheng Liu, Neng-Bo Wang et al. · 0 citations
Review Jul 2026

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier, and synthesizes benchmark design, evaluation protocols, and modeling paradigms based on multimodal alignment, evidence-grounded reasoning, and controlled gener...

Tuo Liang, Zhe Hu, Disheng Liu et al. · 0 citations
Review Open access Aug 2026

Spatial intelligence in vision-language models: a comprehensive survey

This survey provides a comprehensive and unified overview of recent advances in spatial intelligence for VLMs, summarize core concepts behind spatial reasoning in VLMs, analyze why spatial failures occur, and organize existing solutions into a clear framework spanning prompting-based techniques, model improvements, exp...

Di-Sheng Liu, Tuo Liang, Zhe Hu et al. · 9 citations
Preprint Aug 2026

Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models

Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data, is proposed, a training-free framework that recovers a frozen VLA at inference time without policy fine-tuning or failure-specific recovery training.

Yan-Yan Zhang, Di-Sheng Liu, Kai Ye et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.