Skip to content

Author

Jingjing Chen

We have 5 of 169 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

StructRL: Structured Action-Space Exploration for Flow-Based VLAs

Flow-based Vision-Language-Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for adapting them to new tasks. Existing RL methods typically inject stochasticity inside the denoising chain, often through isotropic or temporally independent noise. However, effective robot exploration calls for structured noise: temporally smooth and scaled differently across action groups. We show that simply switching the in-chain noise to a structured form does not suffice: noise added at an intermediate flow time can be weakened by the remaining denoising steps before execution, a phenomenon we call \emph{Structured Noise Dilution}. We propose \textbf{StructRL}, which avoids dilution by relocating policy stochasticity to the action space via three coupled choices: (i) a deterministic ODE decoder, (ii) structured noise injected directly in the action space, and (iii) last-step replay, where policy-gradient updates avoid assigning likelihoods to intermediate denoising states. This keeps structured exploration tied to the executed action while providing a tractable training signal for the flow decoder. Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL. \textbf{Project page:} https://flyfaerss.github.io/structrl/

Jiarui Yang, Bin Zhu, Jingjing Chen et al. · 0 citations
Preprint Jul 2026

DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection

As generative models continue to evolve, AI-generated image detectors must incrementally adapt to emerging generative domains while preserving knowledge acquired from previous ones. This continual learning setting is particularly challenging because forensic traces are often subtle and generator-specific, making detectors highly vulnerable to catastrophic forgetting. Existing methods primarily address this problem by stabilizing feature representations, implicitly treating forgetting as a representation-level issue. In this paper, we show that this perspective is incomplete. We demonstrate that even when feature representations remain discriminative, the decision boundary can progressively drift as the classification head is continually optimized on new domains. These two effects jointly give rise to a compound failure mode, termed Dual Degradation. To overcome this challenge, we propose DECODE, a decoupled continual detection framework that jointly mitigates representation- and decision-level forgetting. Specifically, we introduce Subspace Diversity Regularization (SDR) to preserve diverse forensic representations and Closed-Form Decision Alignment (CDA) to recalibrate the shared classification head after each adapter merge without manual hyperparameter tuning. Extensive experiments on 19 generative domains show that DECODE achieves an average accuracy of 99.36% with only 0.39% forgetting, while further generalizing to 11 unseen generators with 95.36% accuracy.

Zihao Cai, Xinghang Li, Ruiyan Yang et al. · 0 citations
#computer vision Jun 2026

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

It is found that current models often generate visually coherent videos, but struggle with constraint reasoning, counterfactual grounding, physical interaction, and unsafe-instruction suppression, and results show that visual quality and surface-level instruction following are insufficient for trustworthy robotic video world modeling.

Huiqiong Li, Jia-Yu Wang, Zhiting Mei et al. · 4 citations
Preprint Aug 2026

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

Vorch-IR is presented, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model, and an automatic data construction pipeline that synthesizes paired supervision for all four editing settings is developed.

Yaowei Wang, Xiaoyu Chen, Xin Ma et al. · 0 citations
Jul 2026

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention that mitigates this generalized textual bias exerted over visual features that extends beyond isolated sink tokens.

Peng-Kun Jiao, Bin Zhu, Jingjing Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.