Skip to content

Author

Jingjing Chen

We have 9 of 169 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models

Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic fo...

Hui-Qiong Li, Zhi-Ting Mei, Anirudha Majumdar et al. · 1 citation
Preprint Sep 2026

Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement

Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while groundin...

Jia-Rui Yang, Jia-Jin Zhang, Bin Zhu et al. · 0 citations
Preprint Sep 2026

MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

Experiments demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.

Jia-Yu Wang, Bin Zhu, Yue Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition

VICAL is introduced, a consistency-driven framework that improves long-tailed recognition not by enforcing expert diversity, but by reducing prediction variance, and suggests that multi-expert models benefit more from variance reduction than diversity maximization.

Jianggang Zhu, Zheng Wang, Bin Zhu et al. · 0 citations
Preprint Aug 2026

StructRL: Structured Action-Space Exploration for Flow-Based VLAs

Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL.

Jia-Rui Yang, Bin Zhu, Jing-Jing Chen et al. · 1 citation
Jul 2026

DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection

DECODE is proposed, a decoupled continual detection framework that jointly mitigates representation- and decision-level forgetting and introduces Subspace Diversity Regularization to preserve diverse forensic representations and Closed-Form Decision Alignment to recalibrate the shared classification head after each ada...

Zihao Cai, Xing-Hang Li, Ruiyan Yang et al. · 0 citations
#computer vision Jun 2026

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

It is found that current models often generate visually coherent videos, but struggle with constraint reasoning, counterfactual grounding, physical interaction, and unsafe-instruction suppression, and results show that visual quality and surface-level instruction following are insufficient for trustworthy robotic video...

Huiqiong Li, Jia-Yu Wang, Zhiting Mei et al. · 4 citations
Preprint Aug 2026

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

Vorch-IR is presented, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model, and an automatic data construction pipeline that synthesizes paired supervision for all four editing settings is developed.

Yaowei Wang, Xiaoyu Chen, Xin Ma et al. · 0 citations
Jul 2026

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention that mitigates this generalized textual bias exerted over visual features that extends beyond isolated sink tokens.

Peng-Kun Jiao, Bin Zhu, Jingjing Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.